View Full Version : Subtitle Edit
Nikse555
8th October 2011, 12:19
Subtitle Edit 4.0.7 is now out
https://github.com/SubtitleEdit/subtitleedit
SE is an open source (C#) subtitle editor with main focus on creating/editing/sync'ing/adjusting/fixing subtitles, but SE can also import and ocr vobsub and blu-ray image based subtitles (even from matroska/mp4 files), and DVB sub + teletext from .ts files.
SE supports 300+ subtitle formats - let me know if you need more ;)
Can create/edit blu-ray sup and bdn xml files.
Available for Windows and Linux
mastrboy
9th October 2011, 21:43
:) thanks...
Ghitulescu
10th October 2011, 06:59
It looks promising.
I never used it before, that's why I would ask you how well manages SE32 to work with DVD subtitles (SUP), like retiming, synching (to other/preexistent SUP), bitmap editing etc?
Nikse555
10th October 2011, 20:09
It looks promising.
I never used it before, that's why I would ask you how well manages SE32 to work with DVD subtitles (SUP), like retiming, synching (to other/preexistent SUP), bitmap editing etc?
Not too well :(
SE can read and ocr dvds/vobsub/blu-ray sup + a few more image based formats - but the only image based format SE can write is bdn xml/png.
(the blu-ray sup code is converted from 0xdeadbeef's java code for BDSup2Sub)
StainlessS
16th October 2011, 03:36
Crash during "Fix Common Errors".
Crash Report & Mi2.srt here:-
EDIT: Link removed.
Nikse555
16th October 2011, 20:33
Crash during "Fix Common Errors".
Crash Report & Mi2.srt here:-
http://www.mediafire.com/?4z8obiqskr8ikui
Hi StainlessS!
Thx for reporting this :)
Fixed here: http://www.nikse.dk/SubtitleEdit.zip
StainlessS
20th October 2011, 10:46
Thanks Nikse555, got something else to keep you busy.
SubTitle Edit 3.2.2, Build 25663
Crash during Spell Check (HunSpell, dont know if same error as previously reported in other thread)
Crash Report & DWL.srt here:-
http://www.mediafire.com/?6pltpi72a82lz52
MajorX
21st October 2011, 02:19
Thanks Nikse555 for new version of SE.
I have some problem with OCR ...plzz help me...when i use 3.2 OCR of VobSub & Blu-ray sup files are working perfectly but when uninstall it and install new version 3.2.2 my OCR not working now..it only shows orange lines no text. :(
Sample of *.sup subtitles...can u plzz check these subtitles.
http://www.mediafire.com/?zqml8hbrcy6jqgt
Nikse555
22nd October 2011, 07:42
Crash during Spell Check (HunSpell, dont know if same error as previously reported in other thread)
Looks like it's still Hunspell suggest!
I could not re-create this error on my Win7 machine, but I've tried to fix it here (by running suggestions in a separate thread): http://www.nikse.dk/SubtitleEdit.zip
Any better?
Nikse555
22nd October 2011, 18:52
Thanks Nikse555 for new version of SE.
I have some problem with OCR ...plzz help me...when i use 3.2 OCR of VobSub & Blu-ray sup files are working perfectly but when uninstall it and install new version 3.2.2 my OCR not working now..it only shows orange lines no text. :(
Sample of *.sup subtitles...can u plzz check these subtitles.
http://www.mediafire.com/?zqml8hbrcy6jqgt
Aye aye, Major ;)
New version upped: http://www.nikse.dk/SubtitleEdit.zip
Any better?
StainlessS
23rd October 2011, 19:30
Sorry for the delay, Nikse555,
Before I tried your update, I had to download the srt from MediaFire as I did not
keep a verbatim copy. I tried it with the original faulting build 25663, and it did
not fault. Tried this several times, no fault. Got the version srt that I kept,
(probably spell checked via other means) and checked that, same thing, no
fault. Have not ripped any other subs since then (I think) and made no changes
to the setup. I guess it will have to remain a mystery. :confused:
EDIT: Also tried with build 13726, no fault (earlier build No ???).
Chetwood
24th October 2011, 06:51
What do I do to OCR German subs? I've downloaded a German Tesseract package and unpacked it to the program dir but to no avail. I pretty much have to type every word?
Nikse555
24th October 2011, 07:36
Sorry for the delay, Nikse555,
...
Tried this several times, no fault.
...
I guess it will have to remain a mystery. :confused:
Yep, the nhunspell "suggest-method" is not entirely stable
What do I do to OCR German subs? I've downloaded a German Tesseract package and unpacked it to the program dir but to no avail. I pretty much have to type every word?
The German tesseract package should be unpacked to Tesseract\tessdata. Unpacked the file is called deu.traineddata.
Do choose Tesseract as OCR method (not image compare)
And if you're lazy just get this version with German dictionaries included: http://subtitleedit.googlecode.com/files/SE322DE.zip
Chetwood
25th October 2011, 07:55
Mm, I had unpacked it to Subtitle Edit\tesseract\tessdata but ok, your de package is fine, thanks. It also works pretty good, however some events described in parenthesis for the hearing impaired are recognized with mixed case, like
(KEucH†) instead of (KEUCHT)
(I_AcH†) instead of (LACHT).
Also, the small t is recognized as a small l which messes up a lot of items and can only be fixed manually. These new words don't even exit in the German language so shouldn't spellchecking kick in with "prompt for unkown words" being checked? Then the distance between two words ending with r and starting with j is not recognized. Instead of "aber jetzt" it reads "aberjetzt". What can I do to improve this? Thanks.
Nikse555
25th October 2011, 18:04
These new words don't even exit in the German language so shouldn't spellchecking kick in with "prompt for unkown words" being checked?
Yes... problems with loading German dictionary should be fixed here: http://www.nikse.dk/SubtitleEdit.zip
(Tesseract should also be a bit faster now)
Then the distance between two words ending with r and starting with j is not recognized. Instead of "aber jetzt" it reads "aberjetzt". What can I do to improve this? Thanks.
When you press "change all" or "use always" the correction is remembered in OcrFixReplacelist.xml...
xekon
25th October 2011, 22:00
This actually works really good! almost all of the text is right on, and the GUI guides your through smoothly when it needs a fix.
I have a rather strange auto fix though (some kind of error or bug): http://i1208.photobucket.com/albums/cc361/xekon/weirds.png
If you want the .SUP that caused this error to occur give me an email address I can send the file to. (Upon further testing this weird error only happens if the "Try MS MODI OCR for unknown words" checkbox is checked, If I un-check it then this strange substitution does not happen.)
The only OCR error that I get that does not get automatically corrected is the letter "k" being detected as "l<" and not being auto corrected: http://i1208.photobucket.com/albums/cc361/xekon/k.png
I notice similar errors in OCR but they DO get auto corrected like "I\/ly" -> "My"
is there a way I can add "l<" to be autocorrected to "k" ?
also a setting in the options panel to disable "Try MS MODI OCR for unknown words" by default would be handy, then I wouldn't have to uncheck it every subtitle I load
I have also had "d" been detected as "ol" pretty often, then the spell checker dont recognize the word so I edit it manually and change the "ol" to "d"
like the word worried, gets detected as worrieol
Nikse555
26th October 2011, 10:03
...a setting in the options panel to disable "Try MS MODI OCR for unknown words" by default
OK, this setting is now remembered - but I've also improved then check for when to use MODI, so do try to keep it on.
Tesseract (new 3.01 version) now runs in it's own thread, so it should be a bit faster too.
Link to new version:
http://www.nikse.dk/SubtitleEdit.zip
How is it working?
Chetwood
26th October 2011, 14:22
problems with loading German dictionary should be fixed here: http://www.nikse.dk/SubtitleEdit.zip
Mh, this file contains no German dictionary so I copied the one from the de.zip over.
When you press "change all" or "use always" the correction is remembered in OcrFixReplacelist.xml...
The German Umlaut ü (u with two dots above) is often recognized as two i's: ii. Since it's common in several words, how do I replace it for all of them? Thx.
Nikse555
26th October 2011, 14:52
The German Umlaut ü (u with two dots above) is often recognized as two i's: ii. Since it's common in several words, how do I replace it for all of them? Thx.
You can edit [Subtitle edit folder]\Dictionaries\deu_OCRFixReplaceList.xml - add a new WordPart under PartialWords:
<PartialWords>
...
<WordPart from="ii" to="ü" />
</PartialWords>
SE will now look for correct spelled words, where "ii" is replaced with "ü".
You can also take a look at "eng_OCRFixReplaceList.xml".
chainring
26th October 2011, 23:41
Just wanted to stop in here and say thank you for this awesome tool. I love loading up a .sup, letting OCR rip through and having minimal work to correct errors. I can get through an entire movie in 20 minutes.
MajorX
27th October 2011, 03:23
Thanks Nikse555 :)
I have some problem in timing with some subtitles.
when i use OCR...First---I extract subtitles(*.VOBSUB) from video then use it in OCR ..it shows some start time & end time problem like if the original sub have,
Stat Time --> End Time 00:00:13,097 --> 00:00:19,185
OCR shows 00:00:13,097 --> 00:00:17,185
but if i use subtitles direct from video it shows correct start time & end time in OCR.
xekon
28th October 2011, 22:33
I have another feature request, or maybe you know of a configuration file I can edit so that a replacement is always performed.
I would like to replace ’ with ' because ’ shows up very weirdly (last word is supposed to be: didn't ):
http://i1208.photobucket.com/albums/cc361/xekon/didnt.png
Nikse555
29th October 2011, 05:44
@MajorX: This is hard to say why without the actual sub... The ocr window has a check box weather to use time codes from .idx file or from .sub file.
@xekon: Works here in latest version: http://www.nikse.dk/SubtitleEdit.zip
xekon
29th October 2011, 08:01
WOW! you weren't kidding about it going faster! just did a couple more episodes and its zooming through the lines much faster!
edit: odd new bug:
I'm sorry!
was detected as:
I.m
s
0
r
rY
!
http://i1208.photobucket.com/albums/cc361/xekon/imsorry.png
Anakunda
29th October 2011, 15:33
Hello there!
I feel like having trouble with OCR. Recognizing from SUP format, tried both methods and both have significant inaccuracies:
In the pattern comparison mode, the engine totally ignores differencies between letters 'i' and 'l', and 'c' and 'o' and 'e'. All the letters are assigned the character that was assigned by the first occurence of on of letters from "same" group. For example. 1st subtitle contains word more, the wizard stops at o and I assign it o. When it passes over e, it doesnot ask again for letter even if that s 1st "e" in subtitles and assigns it automatically 'o'. That's very bad. I don't know if that's a result of some auto corrections made by SE, but seems to get wrong assigned even if I turn off all the auto corrections on the right side.
That's about character comparison method. Tesseract seems to work better but has considerable flaws too:
Some characters are auto uppercased even if they are in lowercase in the source matrix, especially it concerns 's', 'z', 'c' and 'a'. All occurences of these letters seem uppercased regardless on case in the original matrix if they stand as standalone letter or 1st letter in word. All of s, z, c and a's are kept lowercase if in middle a word.
PPlease give me some suggestions to make functional at least one of the methods, so that most words are recognized properly and don't need to correct by spell checker. The uppercase problem even doesnot seem repairable by spell checker processing!
Thank U !
xekon
29th October 2011, 17:19
I have another feature request, could we have a checkbox to omit all <i> </i> tags, they are being used for only half lines when the whole line is italic, they are also being used when there are no italic lines at all.
Right now after I rip a sub I am going through and doing find/replace to delete them all, but it would be great to have that as a feature in Subtitle Edit.
very often !! gets detected as ll
Is this something that can be fixed? or is there something I can do to help with the detection of exclamation points? or do I have to wait till tesseract is updated?
EDIT: on a side note, whatever you did for MS MODI OCR seems to have worked. and it definitely does help!
here is an example of the ll instead of " or !!
http://i1208.photobucket.com/albums/cc361/xekon/1.png
http://i1208.photobucket.com/albums/cc361/xekon/2.png
xekon
30th October 2011, 09:27
OMG OMG OMG! The programmer in me has just thought of a VERY COOL feature you could add!
call it a visual tool for super fast comparison. (OCR can only get so good, and if you want to verify perfect subs, this is a good way to do it.)
The goal should always be perfect OCR on the first sweep, but visually checking the subs afterwards is just to verify, and the quicker you can do that the better.
Let me know what you think of this idea, I am sure it would actually be something that would be pretty fun to program.
Please let me know what you think because i think it would be AWESOME!
I am drawing an illustration in Photoshop now.
EDIT: ok to illustrate my idea... OCR a .SUP file. then use the arrow key to go down line by line, reading the text, and then looking at the image to compare and see that they are the same.
Now, that is not exactly quick, the brain has to think more, it has to remember more, and your eyes have to move and focus on more than one area, below is my idea:
Basically, use an opengl or directx library that can overlay text, or any library that looks like it will work to overlay text with transparency. And size the text to roughly overlay the SUB image with like a 50-60% transparency. The letters dont have to line up perfectly, anywhere close will allow you to quickly with just a glance tell if the sub and text match visually. (basically you read the sub line ONLY once, and your brain looks for discrepancies as you do it. versus reading two or three times, and moving your eye between locations, and also having to remember and hope you remember correctly.)
I think for somebody that visually checks there OCR for their subs, this would probably speed up the process for them 200%+
see how easy it is to see that they match:
http://i1208.photobucket.com/albums/cc361/xekon/idea.jpg
here is one that passed the OCR, but is incorrect:
http://i1208.photobucket.com/albums/cc361/xekon/wrong1.jpg
here is another one that passed the OCR, but is incorrect (depending on the library used you could even apply a border/stroke to the outside of the letters)
http://i1208.photobucket.com/albums/cc361/xekon/wrong2.jpg
here is another, there is probably one that passes through the ocr, green light and all, in every episode, you just have to look carefully (you might even be able to adjust the thickness of the characters, so that they usually fall within the bounds of the SUB image character outlines):
http://i1208.photobucket.com/albums/cc361/xekon/wrong3.jpg
Chetwood
31st October 2011, 07:56
Looks impressive but I think it's overkill. Why not simly have a small window showing the item and an editable text window below that shows the OCRed text. In case they don't match simply alter the text and move on to the next item.
Nikse555
31st October 2011, 19:20
Thanks Nikse555 :)
I have some problem in timing with some subtitles.
when i use OCR...First---I extract subtitles(*.VOBSUB) from video then use it in OCR ..it shows some start time & end time problem like if the original sub have,
Stat Time --> End Time 00:00:13,097 --> 00:00:19,185
OCR shows 00:00:13,097 --> 00:00:17,185
but if i use subtitles direct from video it shows correct start time & end time in OCR.
My guess would be that the application you ripped the vobsub with did not use time codes from the mkv container, but rather used the time codes in the sub file itself (the time codes in idx and sub file are exactly alike).
xekon
31st October 2011, 19:43
Nikse555 please let me know what you think of my idea, if its not something your interested in, then I will try adding it. I just noticed Subtitle Edit is open source.
Could I please have a copy of the source code that is as current as: http://www.nikse.dk/SubtitleEdit.zip
the one on code.google.com is October 14.
xekon
31st October 2011, 19:47
Looks impressive but I think it's overkill. Why not simly have a small window showing the item and an editable text window below that shows the OCRed text. In case they don't match simply alter the text and move on to the next item.
Subtitle Edit has very accurate result for the OCR. There are usually only 1-3 wrong subs out of 300 lines. That is quite impressive. So generally you wont need to do much editing, only verifying. The method I posted is the quickest way that I can think of to scan through entire sub files after the OCR and visually verify.
Nikse555
31st October 2011, 20:49
Hi Anakunda!
...
In the pattern comparison mode, the engine totally ignores differencies between letters 'i' and 'l', and 'c' and 'o' and 'e'. All the letters are assigned the character that was assigned by the first occurence of on of letters from "same" group. For example. 1st subtitle contains word more, the wizard stops at o and I assign it o. When it passes over e, it doesnot ask again for letter even if that s 1st "e" in subtitles and assigns it automatically 'o'. That's very bad.
Yes, this is true. I've tried to improve it a bit here: http://www.nikse.dk/SubtitleEdit.zip
A work-around is to right-click on the offending line in the list view, and choose "Inspect compare matches for current image" - here you can choose "Add better match" to correct mistakes.
(my image compare code is a bit slow for blu-ray images...)
Tesseract seems to work better but has considerable flaws too:
Some characters are auto uppercased even if they are in lowercase in the source matrix, especially it concerns 's', 'z', 'c' and 'a'. All occurences of these letters seem uppercased regardless on case in the original matrix if they stand as standalone letter or 1st letter in word. All of s, z, c and a's are kept lowercase if in middle a word.
Is this still the case in latest version?
If yes, could you provide a test file + a few line numbers?
Nikse555
1st November 2011, 08:53
...
call it a visual tool for super fast comparison. (OCR can only get so good, and if you want to verify perfect subs, this is a good way to do it.)
...
Another way to proof read would be to right click in the list view - and choose "Save all images with html index...". This displays a web page with all images + ocr'ed text if available. In latest version, this also shows text with background color.
MajorX
2nd November 2011, 03:03
Hi Nikse555
can u plzz check this *.SUP file...i get only strange symbols with OCR.
http://www.mediafire.com/?aoy66c5ue9mbah9
xekon
2nd November 2011, 03:08
MajorX I tried your file with Nikse555's latest version here: http://www.nikse.dk/SubtitleEdit.zip
I also got lots of symbols if I had "Try MS MODI OCR for unknown words" unchecked.
but if you use the MS MODI OCR it detects all of them just fine :)
give it a shot.
PS: I wonder if that subtitle file has ever had its resolution resized.... the letters are really bad quality.
MajorX
2nd November 2011, 07:36
I try with this version but i can't enable MS MODI OCR ...can u tell how can i do this.
http://img266.imageshack.us/img266/1182/74087558.png
kypec
2nd November 2011, 10:50
I try with this version but i can't enable MS MODI OCR ...can u tell how can i do this.
You must have some Microsoft Office libraries installed for this to work IIRC...
Nikse555
2nd November 2011, 22:30
Hi Nikse555
can u plzz check this *.SUP file...i get only strange symbols with OCR.
Thx for the file :)
This font don't look blu-ray like but seems clear enough. Resizing did not help, but changing font color to white seems to help, so this is included latest version, which should handle your sup better: http://www.nikse.dk/SubtitleEdit.zip
MajorX
3rd November 2011, 06:06
Thx for the file :)
This font don't look blu-ray like but seems clear enough. Resizing did not help, but changing font color to white seems to help, so this is included latest version, which should handle your sup better: http://www.nikse.dk/SubtitleEdit.zip
Thanks Nikse555....working perfectly. :) :)
Nikse555
13th January 2012, 15:17
Subtitle Edit 3.2.3 is now finally out with lots of minor improvements and fixes!
Change log
New: Added Brazilian Portuguese - thx XXXXXXXXXX
New: Added Italian language file - thx Maff
New: Added Portuguese (Portugal) language file - thx Ricardo Perdigão
New: Added Japanese language file - thx Nardog
New: Added Spanish language file - thx m2s
New: Support for subtitle format AvidCaption - thx Laszlo
New: Support for F4 subtitle formats - thx Fred
New: Export to Blu-ray sup format
Improved: Updated Tesseract to 3.01. Now includes (some) italic detection + adds support for Arabic, Hebrew, Hindi and Thai
Improved: Undo improved so it also works for textbox + redo (Ctrl+Y)
Improved: Many new configurable shortcuts (e.g. for fullscreen video player)
Improved: OCR tweaked a bit + BluRay sup files are processed faster
Improved: TextBox with current subtitle now shows cursor position - thx Leszek
Improved: Subtitle format PAC much improved - thx Peter
Improved: Subtitle format FCP Xml improved - thx Ulrik
Improved: Subtitle format D-Cinema improved - thx Karam
Improved: Splitting of lines - Thx Trottel
Improved: Auto break lines - thx Majid
Improved: Some fixes for Fix common errors/Remove text for HI - thx Majid
Improved: Optimized Fix Common Errors
Improved: DirectShow can now also play audio-only files
Fixed: Crash when setting Options - thx karmazyn
Fixed: Crash in set color (or set font) - thx LEO33
Fixed: Crash/freeze when loading large subtitle files - thx Ulrik
Fixed: Bug when clicking in list view while running ocr - thx sialivi
Fixed: De-selecting text in textbox via single click - thx XhmikosR
Fixed: Possible crash in spell check + German dictionary should work
Fixed: Missing save/load of a fix common errors setting - thx menes
Fixed: Removed Microsoft translate as it's useless with new quotas
Fixed: Milliseconds in timed text - thx Calle
Fixed: Names with spaces now works in spell check - thx Dr. jackson
Fixed: Do not use frame rate if it's zero (audio files) - thx dixie.fever
Fixed: Possible crash when saving xml files - thx Peter
http://code.google.com/p/subtitleedit/downloads/list
kalehrl
13th January 2012, 16:01
Thank you Nikse.
This is the best and most complete subtitle editor ever.
tonyymmao
22nd January 2012, 01:06
this really is good, i just have one question though, is it possible to add a border option for the text, cuz i've seen some videos with image based subs have different colour borders like ass.
Nikse555
22nd January 2012, 09:46
this really is good, i just have one question though, is it possible to add a border option for the text, cuz i've seen some videos with image based subs have different colour borders like ass.
For which subtitle format?
SE allows for styles for in .ass files, but otherwise SE does not offer other styles than italic, bold, and font color in SubRip/MicroDvd files.
Latest SE test version should be able to export VobSub (and bd sup with correct timestamps) via File -> Export.
Could anybody verify that it works?
http://www.nikse.dk/SubtitleEdit.zip (or get the C# code from svn and compile it yourself)
Chetwood
23rd January 2012, 14:52
I've set SE to ANSI as the default encoding type, OCR'ed a Vobsub and saved it to SRT. When I open it again in SE with "Autodetect ANSI encoding" checked, all German Umlauts are messed up. When unchecked the umlauts are ok.
When exporting to VobSub on default settings the sub is far less readable than the original VobSub. Gonna pm you some files.
OtonVM
4th April 2013, 09:59
This is an amazing piece of software, thank you for making it!
I have found a slight problem with FAB subtitles in Encore.
I export from srt to fab so I can keep italics and such. What I get is an image like this (tinypic converted from tiff to jpeg removing the alpha channel so black is actually transparent):
http://i45.tinypic.com/4u8ggj.jpg
with it's script:
IMAGE002.tiff 00:03:01:01 00:03:03:18 123 417 596 465
and this is the result:
http://i48.tinypic.com/2upwrnp.png
notice the line above the text.
What I have to do is lower (I think it's lower) the image by 1px:
IMAGE002.tiff 00:03:01:01 00:03:03:18 123 416 596 465
so I can get this:
http://i48.tinypic.com/2dig4g2.png
One line is not a problem but it happens multiple times, seemingly at random.
Nikse555
8th May 2013, 08:04
Latest version of SE (installer version only) uses the roaming profile. On my computer it's C:\Users\Nikse\AppData\Roaming\Subtitle Edit\Tesseract\tessdata
If you edit your Settings.xml file (C:\Users\<username>\AppData\Roaming\Subtitle Edit\Settings.xml) and change <ShowBetaStuff>False</ShowBetaStuff> to <ShowBetaStuff>True</ShowBetaStuff> you will get a "..." button that can download/install Tesseract languages for you in the ocr window.
@OtonVM: Sorry about not replying sooner - do you still need this?
OtonVM
8th May 2013, 08:31
Not a problem, I was busy too... :)
But yes, I did not solve this problem. I created a python script that modified the txt file but it did not help.
Betsy25
17th May 2013, 00:18
Regarding the "fix common errors..." command, is there some way we can configure this ? For instance change the default "Break long lines" from 45 to some other value ? :stupid:
Nikse555
18th May 2013, 22:14
Try Options -> Settings -> General - Single line max. length (I think...)
Betsy25
21st May 2013, 21:44
Try Options -> Settings -> General - Single line max. length (I think...)
Doh. Thank you very much Nikse.
Betsy25
25th May 2013, 17:53
Just have a weird issue here. I use the "default" Directshow video engine as the video player, which also auto-adds the subs into the display on it's own (DirectVobSub), but when playing *any* mkv video the subs drawn by SE are almost 0.2 secs late compared to the default renderer's (which are by DirectVobSub, and I suspect are the correct timings).
Please, why is there such a relatively large delay between what DirectVobSub renders (and which will be the timings we get when we play our content in our players of choice) and the ones that get drawn on screen by SE ?
Seeing this, I simply cannot trust altering the timings using SE that way :confused:
JeanMarc
26th May 2013, 20:19
Thank you for this high quality subtitle application.
When I want to convert this sup subtitle file (extracted via ea3cto.exe from a blu-ray file) to subrip, it works very well in the normal GUI mode.
If I want to do the same thing in command line mode, for example:
SubtitleEdit /convert cotgf.sup subrip, I get the following output:
E:\Cotgf_test>
Subtitle Edit 3.3.3 - Batch converter
- Syntax: SubtitleEdit /convert <pattern> <name-of-format-without-spaces> [/offs
et:hh:mm:ss:msec] [/encoding:<encoding name>] [/fps:<frame rate>] [/inputfolder:
<input folder>] [/outputfolder:<output folder>]
example: SubtitleEdit /convert *.srt sami
list available formats: SubtitleEdit /convert /list
1: E:\Public\Public_videos\_videos to do (HD)\Cotgf_test\cotgf.sup - input file
too large!
I did'nt find any reference to that issue anywhere.
I am using Subtitle Edit 3.3.3 rev.1745 on Windows 7, and the sup file is 22,537KB.
Thank you for any help.
Nikse555
7th June 2013, 15:01
@OtonVM: What happens if you use full resolution at each line (the last two parameters), like:
IMAGE002.tiff 00:03:01:01 00:03:03:18 123 417 720 480
@Betsy25: Argh, I only checked for subtitles every 250 ms :thanks:
Should be fixed in this beta: http://www.nikse.dk/SubtitleEdit.zip (portable version)
@JeanMarc:
I have not added ocr to command line... only to batch convert.
Do you need to command line version?
Also, I still recommend to run ocr live, so error checking is easier.
Betsy25
7th June 2013, 18:39
...
@Betsy25: Argh, I only checked for subtitles every 250 ms :thanks:
Should be fixed in this beta: http://www.nikse.dk/SubtitleEdit.zip (portable version)
...
Thanks, all good ! :o
Chetwood
12th June 2013, 06:01
I was unaware this tool can OCR from SUPs inside MKVs which is nice. It would be very cool to be able to do this from m2ts as well so there'd be no need to extract subs to find forced items not properly marked as such prior to conversion. Thx.
THXTEX
16th June 2013, 22:54
Just tested this awesome tool. Looks like it can do a lot of stuff. I used the export to BD .sup, but there seems to a be a little bug in that function. If there is a p, g, q or j - letters that has a part below the baseline that subtitle is placed higher than subtitles not having this.
Anyway - a great tool. Thx
mariner
31st July 2013, 10:25
Greetings Nik. Many thanks for this wonderful tool.
Appreciate if you could kindly explain the correct way to edit the time code of BD sup and save the modifications. It seems all png information is lost in the process.
Best regards.
.....
<Event InTC="00:01:31:00" OutTC="00:01:32:11" Forced="False">
<Graphic Width="150" Height="50" X="885" Y="1015">0002.png</Graphic>
</Event>
<Event InTC="00:01:37:06" OutTC="00:01:39:06" Forced="False">
<Graphic Width="150" Height="50" X="885" Y="1015">0003.png</Graphic>
</Event>
.......
THXTEX
24th August 2013, 09:20
Just tested this awesome tool. Looks like it can do a lot of stuff. I used the export to BD .sup, but there seems to a be a little bug in that function. If there is a p, g, q or j - letters that has a part below the baseline that subtitle is placed higher than subtitles not having this.
Anyway - a great tool. Thx
There are still some issues regarding baseline alignment. When there is a , (comma) the line still floats above the baseline. Other lines are also moving differently up but I can not see any pattern there.
Playing back a m2ts muxed with tsMuxer and subtitle from SE337 in VLC does not work correctly. The subtitle is only shown for a fraction of time. Playing back the same file in Media Player Classic is OK.
Nikse555
30th August 2013, 20:15
There are still some issues regarding baseline alignment. When there is a , (comma) the line still floats above the baseline. Other lines are also moving differently up but I can not see any pattern there.
In latest svn I've tried to improve this a bit... or here: http://www.nikse.dk/SubtitleEdit.zip (portable, beta)
Playing back a m2ts muxed with tsMuxer and subtitle from SE337 in VLC does not work correctly. The subtitle is only shown for a fraction of time. Playing back the same file in Media Player Classic is OK.
With Bluray sup?
THXTEX
1st September 2013, 07:42
In latest svn I've tried to improve this a bit... or here: http://www.nikse.dk/SubtitleEdit.zip (portable, beta)
Subs still moves up and below the baseline - but's getting better.
With Bluray sup?
Yes
mariner
1st September 2013, 15:52
Greetings Nik. Many thanks for this wonderful tool.
Appreciate if you could kindly explain the correct way to edit the time code of BD sup and save the modifications. It seems all png information is lost in the process.
Best regards.
.....
<Event InTC="00:01:31:00" OutTC="00:01:32:11" Forced="False">
<Graphic Width="150" Height="50" X="885" Y="1015">0002.png</Graphic>
</Event>
<Event InTC="00:01:37:06" OutTC="00:01:39:06" Forced="False">
<Graphic Width="150" Height="50" X="885" Y="1015">0003.png</Graphic>
</Event>
.......
Anyone?
loninappleton
1st September 2013, 21:21
Some responses on here begin in 2011 and then jump to the current day all in 4 Doom9 pages.
Please give a source for a guide for use of the current version of Subtitle Edit.
I have been struggling with an older film in French that has only partial English subtitle srt file. There is a prologue which is not translated (any help appreciated, details can be mailed) and so dialog starts in at about 2 minutes after the beginning. I had done a timing on it at 2:38 IIRC.
From there I think there is an srt sub for NTSC only available for one extant PAL download or vice versa.
I can get the details from GSpot if someone picks up on this query. Simply put, I have never been able to get the voice subs lined up no matter what technique I have tried to stumble through. Perhaps there is a way to deconstruct these errors?
One thing I know is the dialog is always in advance of the video and I have not been able to discern whether timing in consistent through out.
This is a pet project. The film is of minor interest to a minimal audience. Fans of directors like Jean Rollin and Jess Franco may find it interesting. Others may find the exploitative nature of the content (pretty mild) off-putting and so my advisory is given.
Nikse555
2nd September 2013, 20:56
@mariner: Sorry, SE cannot edit bd sup files - only import and export atm.
@THXTEX: Hm, I can see that it also happens with bdsup2sub but not with bdsup2sub++! Anybody knows about this? Well, I guess you can re-save the sup file with bdsup2sub++ to make it work.
@loninappleton: If you want to add subtitles to an existing srt file, I would use the waveform and 'draw' the subtitles (or setup some shortcuts via Options -> Settings -> Shortcuts). The SE help file is here: http://www.nikse.dk/SubtitleEdit/Help#video_modes
johner23
2nd September 2013, 23:07
Hello, dear all!
@Nikse555
SE cannot edit bd sup files - only import and export atm.
Nowadays, which programs can edit bd sup (and similar ones) files accordingly?
In a future, maybe SE can do it also?
Thanks.
Best regards.
mariner
3rd September 2013, 06:33
@mariner: Sorry, SE cannot edit bd sup files - only import and export atm.
...
Thanks for the kind reply, Nikse555.
At the moment, I can get SE to read a XML file (which is just a text file) and modify the time code, but have problem saving the changes made. Is it possible to get SE just to save the modified time codes while leaving the other parameters of the XML file untouched?
Appreciate if you'd kindly look into this.
Many thanks and best regards.
mariner
3rd September 2013, 06:40
Hello, dear all!
@Nikse555
Nowadays, which programs can edit bd sup (and similar ones) files accordingly?
In a future, maybe SE can do it also?
Thanks.
Best regards.
Try BDSup2Sub.
THXTEX
4th September 2013, 20:45
Well, I guess you can re-save the sup file with bdsup2sub++ to make it work.
Yes, that works. But the re-saved sub is about 33% larger in filesize....Strange.
Still the baseline alingment issue.
Nikse555
5th September 2013, 13:53
Yes, that works. But the re-saved sub is about 33% larger in filesize....Strange.
Still the baseline alingment issue.
OK, the short display time might be fixed now: http://www.nikse.dk/SubtitleEdit.zip
About the baseline alignment... it's pretty close on my computer. Do you have some test cases where it's bad?
THXTEX
5th September 2013, 21:50
Do you have some test cases where it's bad?
I have pm'ed you an example. Select the Danish sub.
You'll notice that subs with a part below the baseline are placed a bit higher than lines without. But at the end the lines with part below baseline are placed lower.
Nikse555
6th September 2013, 17:04
I have pm'ed you an example. Select the Danish sub.
You'll notice that subs with a part below the baseline are placed a bit higher than lines without. But at the end the lines with part below baseline are placed lower.
New version upped: http://www.nikse.dk/SubtitleEdit.zip
Baseline alignment should be fixed :)
THXTEX
7th September 2013, 08:44
Think I have found what is causing the strange behaviour. I have PM'ed you a link for a new testfile. The file has 2 subs - both created with SE. In one of them I used Ariel (50) font and on the other Verdana (50) font. In Verdana the letter 'å' has a large circle which makes letter very tall. This somehow pushes the entire line below the baseline. The sub created with Ariel looks fine. Does this make sense?
Anyway I will go with the Ariel font which anyway looks better.
..... and then some more suggestions for the SUP export function. :-) Add Shadow option and fade. The Shadow could be in 8 directions. Regarding fade i'm not sure how it's added but have seen it in another software.
kevmitch
18th September 2013, 03:47
Playing around with Subtitle Edit 3.3.8 rev.2047 on Linux. My source material are the subtitles from the Criterion blu-ray release of Solaris. There are lots of itallics! Attached are a non-italic subtitle (Image53) and italic subtitle (Image54).
It seeems that the un-itallic factor (right click on an individial title in the ocr import list) is getting automagically set to what appears by eye to be the correct value. The resulting preview text looks as non-itallic as it gets. However, Tesseract isn't having any of it. It completely barfs all over the italics although it handles the non-italic stuff fine. The exported italic subtitle (Image54) is the original, while I can't find a way to export an unitalicized image.
I get the same bad OCR results running Tesseract 3.02.02 directly on this Image54. This suggests that Subtitle Edit is handing Tesseract the original italicized image, and not the unitalicized one. Is this the case?
infante
4th October 2013, 23:00
Hi All,
Anyway to remove all text markup tags (e.g. <i> Text </i>) without going through line by line?
thanks!
kalehrl
5th October 2013, 07:40
CTRL+A to select all, then right click and choose Normal.
osgZach
9th November 2013, 22:52
Maybe I'm being incredibly dense. But is it possible to re-open the OCR window?
i.e I start working on something, but then I want to save my work and come back to it later... ?
Nikse555
10th November 2013, 01:04
It's not really super easy to see... but if you've saved a partial ocr'ed file, you can open the image based subtitle again and right-click in the list view and choose 'Import text with matching time codes...'
kalehrl
10th November 2013, 09:01
Is it possible to add an option in the program to split 2 long lines of subtitles into 3 lines?
My dvd player can only show 42 characters in a single line.
Sometimes I have 45 characters in one line and 50 in another.
The program would then split them into 3 lines which would then be shown appropriately by setting max line length to 42 in the options.
Thank you.
osgZach
11th November 2013, 00:51
It's not really super easy to see... but if you've saved a partial ocr'ed file, you can open the image based subtitle again and right-click in the list view and choose 'Import text with matching time codes...'
That's really Strange.. usually what I am doing is I OCR with Subtitle Edit, then save those results after proofing as a .ASS file and do any further typsetting in Aegisub..
So I re-opened an SUP, and tried to import, but only some lines show up.. but I haven't modified their timings at all when working on them :rolleyes: I wonder if Aegisub is truncating to tens of milliseconds O_o
minhjirachi
11th November 2013, 04:27
Please add the shadow option when export to.sup file like easySup. Waiting for this function
Nikse555
17th November 2013, 15:34
@kevmitch: Yeah, I get the same behaviour, but I cannot debug as Ubuntu/Mono Develop will not run for me...
@kalehrl: SE cannot auto break lines to three lines atm, but you can try "Tools" -> "Split long lines" instead.
@osgZach: Thx for reporting this re-open ocr issue - time codes in ass/ssa is less precise, so it did not work too well. Will be fixed in next version (or get latest beta or latest src).
@minhjirachi: Shadow option when exporting to blu-ray sup or bdn/xml is available in latest beta/latest src.
Latest beta: http://www.nikse.dk/SubtitleEdit.zip
lansing
24th November 2013, 22:12
can you also add the tesseract ocr dictionary for chinese traditional?
lansing
25th November 2013, 06:05
I'm trying to ocr the chinese characters through the image compare method, but the characters that were being detected were mostly cropped on all sides. Is there anyway to expand the crop on all sides?
And also, I wanted to request adding custom hotkey for expand and shrink selection in the ocr window, because it will save a lot of time while doing heavy ocr without using mouse and keyboard at the same time.
Nikse555
25th November 2013, 22:30
can you also add the tesseract ocr dictionary for chinese traditional?
Sure, will be there in next version - or you can get this: http://www.nikse.dk/SubtitleEdit.zip
And also, I wanted to request adding custom hotkey for expand and shrink selection in the ocr window...
Try Alt+Arrow left/right
No way to expand the crop atm, sorry.
lansing
27th November 2013, 08:54
thanks for the reply, alt+arrow left/right works. Now I'll try to ocr an small image database to see it if it's accurate even with the cropped sides
LowDead
12th December 2013, 20:25
Hi! Thanks for a wonderful piece of software.
I'm having trouble syncing and wondered if there was something I missed. In SE 3.3.11 using vlc and visual sync, I get perfect sync. Then saving the srt and muxing it with the blu-ray the sub gets gradually out of sync. With Directshow i don't get audio just video but it seems that the speed of the movie is somewhat different(maybe the correct speed to sync to?) but when looking at the little info text just above the waveform window it says 23.976 both when using vlc and directshow. Any ideas?
//LD
Never mind, it was bad muxing on my side.. :o
minhjirachi
20th December 2013, 02:12
I have a little problem with subtitle edit. When I export to .sup file, with this line: "@facebook.com/VAV.AudioViet", the text: "@facebook.com" has no shadow and "VAV.Studio" is normall.
<font color="#ff8000">V</font><font color="#008000">AV</font> <font color="#ffffff">STUDIO</font> .:. <font color="#ff8000">V</font><font color="#008000">AV</font>.MOVIE
©facebook.com/<font color="#ff8000">V</font><font color="#008000">AV</font>.AudioViet
And with that script, the subtitle edit won't align center. I don't know why. So please fix that problem.
Thank you.
Nikse555
20th December 2013, 09:54
I have a little problem with subtitle edit. When I export to .sup file, with this line: "@facebook.com/VAV.AudioViet", the text: "@facebook.com" has no shadow and "VAV.Studio" is normall.
Thx for reporting this :)
Will be fixed in next update (you can also get latest source or latest beta: http://www.nikse.dk/SubtitleEdit.zip)
Centering should work if you use correct video size - as far as I know.
minhjirachi
22nd December 2013, 03:52
Thx for reporting this :)
Will be fixed in next update (you can also get latest source or latest beta: http://www.nikse.dk/SubtitleEdit.zip)
Centering should work if you use correct video size - as far as I know.
Thank you so much. The beta was fixed the shadow problem but the centering still get bug.
Waiting for your new release.
minhjirachi
10th January 2014, 14:56
I have another problem. When I export the .sup file and remux it with tsmuxer. After that I demux it with BD Reauthor Pro and allow it auto generate the project. I can't mux it with that Subtitle Edit sup file. I have to use the BDSup2Sub to export the subtitle as xml/png and re-import into Scenarist as BDN. So please fix this problem too and add the same features "bottom margin" of "Export Blu-ray sup" to "Export BDN xml/png".
Thank you so much.
Waiting for your new official release.
Nikse555
10th January 2014, 21:49
OK, latest src now have "Bottom margin" option for bdn/xml.
I just tested tsMuxeR and SE sup files and it seems to work... so I have no clue about what to fix or how to test :confused:
Next version of SE will be out in a few days (also with some support for ripping image based DVB subtitles)
Nikse555
12th January 2014, 15:47
SE 3.3.12 final release out: http://code.google.com/p/subtitleedit/downloads/list
I hope that SE can now import image based DVB SUB from .ts files... please do let me know if you have problems!
If you want to OCR and the sub contains colored (non-white) lines, then do check the 'Grayscale' checkbox.
Dean007
13th January 2014, 20:39
Hello! I'm new to this software. :)
I tried ripping Slovenian subtitles from a DVD (pal). There were a few issues. I noticed the software likes to put different letters or words by itself. For instance, word NO... was ripped as M
I have "try to guess word" unchecked and "prompt for unknown words" checked.
And many words were typed together (for instance; what if --> whatif, etc.), also a lot of times letters would begin with upper letters when there's no ?!.
I downloaded slovenian dictionary and tessarect and copied them in right folder.
Any suggestions how can I improve things? And can I edit OCR fix?
Nikse555
14th January 2014, 17:50
Hi Dean007, OCR via Tesseract relies on the Tesseract language data files, so SE cannot make a big differences in the results - only some post OCR fixes. Check the file for English post corrections to see how it's done: eng_OCRFixReplaceList.xml.
It's possible to train Tesseract yourself, but it's NOT for everybody... at all: http://code.google.com/p/tesseract-ocr/wiki/TrainingTesseract3
Tesseract is normally very good for OCR'ing English, but you might want to try the OCR via "Image compare" in SE.
Dean007
14th January 2014, 19:19
I see. :(
With image compare, does SE recognizes typed letters or do I have to keep manually type same letters again and again?
Nikse555
15th January 2014, 07:34
With image compare, does SE recognizes typed letters or do I have to keep manually type same letters again and again?
Depends on the image quality and your 'Max. error%' (should be around 1.0 for vobsub like subs and probably around 5-7 for larger subs like bluray).
Unless you have multiple fonts in the same sub or uses resized vobsub files it should work fine...
lansing
15th January 2014, 11:09
is there a selection option to select even/odd lines?
Nikse555
15th January 2014, 20:05
is there a selection option to select even/odd lines?
No, is that a commonly used functionality for you?
lansing
15th January 2014, 20:58
No, is that a commonly used functionality for you?
It's for the subtitle that I'm working on, the original person who's doing the timing screwed up on the import, and now half of the file were junks under this pattern. So I'm looking for a way to remove them.
lansing
17th January 2014, 03:34
I encountered a bug while editing, I have my subtitle loaded in, working with the waveform for syncing, and saving constantly. However, after a while, when I try to save it, I noticed that it was saving it to a temp file called "untitled" and the program asks me if I want to overwrite it. I don't know what triggered the bug.
And I think the scrolling direction in the waveform is wrong, down scroll should advance the time in the waveform, not the opposite.
Nozdrum
17th January 2014, 17:40
By any chance, is it possible to OCR an hard subbed .mp4 file using Subtitle Edit? this mp4 has no special characters, it uses one font for the whole duration, it looks like they just hard subbed the .srt track but I cannot find the soft subbed version of it so I wonder if there's any way to OCR it, any suggestions?
Nikse555
17th January 2014, 21:29
@lansing:
I just might have added this select even/uneven lines in Edit -> Modify selection.
The "Untitled" issue might have been fixed too... I think it was related to the undo-function.
The waveform mousewheel scrolling direction can be changed in Options -> Settings -> Waveform/spectrogram
http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
How does it work?
@Nozdrum: You can try to open the .mp4 file in SE via File -> Open - to check if it contains subtitle tracks (vobsub/bluray sup).
SE cannot extract hardcoded/burned-in subtitles, sorry.
lansing
18th January 2014, 09:04
the select even/odd function work flawlessly, thanks for adding it.
I have another request, that is to add a hover-to-focus to the waveform, because I'll have to constantly switching back and forth between the list view and waveform window while syncing subtitle. Now I have to click somewhere inside the waveform in order to have focus on it, but doing so will move the current position of the waveform cursor, which is not good. And it's going to add several hundreds of unnecessary clicks to the workflow.
Nikse555
18th January 2014, 10:13
@lansing: nice idea about the waveform auto-focus on mouse over :) - I've added it in latest svn and here: http://www.nikse.dk/SubtitleEdit.zip
lansing
18th January 2014, 11:01
@lansing: nice idea about the waveform auto-focus on mouse over :) - I've added it in latest svn and here: http://www.nikse.dk/SubtitleEdit.zip
Doesn't quite work yet, it still set focus on the waveform even though I hovered out.
Nikse555
18th January 2014, 14:18
@lansing: Focus something else than current focused control on mouse out/lease sounds complicated/confusing - what should be focused and how to keep focus?
Instead I've made two custom shortcuts: One for shifting focus from list view to waveform + one for shifting focus from waveform to list view... anyone got a better idea?
Test version: http://www.nikse.dk/SubtitleEdit.zip
lansing
18th January 2014, 20:03
@lansing: Focus something else than current focused control on mouse out/lease sounds complicated/confusing - what should be focused and how to keep focus?
Instead I've made two custom shortcuts: One for shifting focus from list view to waveform + one for shifting focus from waveform to list view... anyone got a better idea?
Test version: http://www.nikse.dk/SubtitleEdit.zip
I'm thinking of switching the focus depending on the view window, as they are the main areas that can use mouse scrolling.
So if the current view window is "list view", focus will be set between waveform and list view window when mouse hover: hover over waveform->waveform get focused, hovered out waveform->list view get focused. And if "source view" window is selected, the focus switching will be between waveform and source view window instead.
Betsy25
20th January 2014, 21:11
Hi,
The Hunspell spelling engine is quite clunky at finding words, and things like 'Forrest' pops up pointing that Forrest' is not a recognized word (happens with the dictionary set to Dutch), and all kind of clumsy "mistakes" like 1 letter words.
Wouldn't it be too hard to implement a much smarter engine like Aspell please ?
Nikse555
20th January 2014, 21:45
@lansing: OK, I've added an option to also focus the "list view" on mouse enter: http://www.nikse.dk/SubtitleEdit.zip
@Betsy25: Prompt for unknown "One letter words" is an option in Settings -> Tools. Perhaps you can find a better Dutch dictionary? Hunspell seems to be used more than Aspell... Yet another option is the "Word spell check" plugin.
minhjirachi
21st January 2014, 04:55
Still get center alignment bug. Please fix it.
Nikse555
21st January 2014, 19:46
Still get center alignment bug. Please fix it.
How is this version: http://www.nikse.dk/SubtitleEdit.zip ?
If it still has center alignment issues, then which text + font causes the problem?
minhjirachi
23rd January 2014, 15:19
I have 3 scripts cause centering problem.
First script:
<font color="#FF0000">VAV STUDIO Hân hạnh Giới thiệu bộ phim:
<b><i>ĐỊCH NHÂN KIỆT</i></b>
</font>
It's cause problem on version 3.3.12 and no more problem on 3.3.13.
Second script and third script:
<font color="#FFFF80"><b><i>Phụ đề Việt ngữ bởi
trwng_tamphong ~ phudeviet.org</i></b>
</font>
<font color="#FFFF80">Phim thuyết minh độc quyền tại
<b><i>www.thuyetminh.net</i></b>
</font>
Still happend on the latest version of Subtitle Edit.
Betsy25
23rd January 2014, 16:27
@Betsy25: Prompt for unknown "One letter words" is an option in Settings -> Tools. Perhaps you can find a better Dutch dictionary? Hunspell seems to be used more than Aspell... Yet another option is the "Word spell check" plugin.
Yeah, I know that option, the problem is that the dutch language as loads of occurences with a ' in front of it, like 't for het (it), 's for eens (once), 'n for one (a), etc....
and the Hunspell engine doesn't seem to get it, so does the word check plugin, with that option unchecked it still halts at every occurence of such words :(
Nikse555
23rd January 2014, 18:08
@minhjirachi: thx for the examples :)
Should work now: http://www.nikse.dk/SubtitleEdit.zip
@Betsy25: Hm, it might actually be my code that prevented these 't 's and 'n from working... is this the above version any better?
Also, there's a slightly newer spell check on this page I think: http://www.opentaal.org/bestanden/doc_download/20-woordenlijst-v-210g-voor-openofficeorg-3 (rename oxt to zip and copy .aff and .dic files to SE\Dictionaries)
Ghitulescu
23rd January 2014, 18:56
Any progress with the extraction of subtitles from TS/M2TS streams? In particular from HD broadcasts?
Nikse555
23rd January 2014, 19:06
Any progress with the extraction of subtitles from TS/M2TS streams? In particular from HD broadcasts?
The above version should be able to open and extract bitmap based subtitles from .ts files... and it might work with .m2ts files as well (just use File -> Open).
Please test :)
Betsy25
23rd January 2014, 21:12
@Betsy25: Hm, it might actually be my code that prevented these 't 's and 'n from working... is this the above version any better?
Also, there's a slightly newer spell check on this page I think: http://www.opentaal.org/bestanden/doc_download/20-woordenlijst-v-210g-voor-openofficeorg-3 (rename oxt to zip and copy .aff and .dic files to SE\Dictionaries)
Hi Nikse, using your updated version, way better but still some other instances like 'm for hem (him), 'r for haar (her), 'k for ik (I), 'n is still there.
in the "Word not found" textfield it asks what to do with m etc... (without the antecedent ' ), I can't quite choose to add to user database because that might lead to some really crappy user database quite fast , it might be correct if it presented to add the precedent ' AND the letter together to the database though. (but then it probably would be worse for the case if the language was English, perhaps)
now, the problem pointed out before, now popping up for replacement of occurrences like 'Forrest' not recognized, there is nothing painted in red, and the "Word not found" textfield asks what to do with Forrest' again.
BTW, Replacing the dutch dictionary didn't change anything to the problems I'm having.
Duh, What I found, I can of course add all dutch '{one letter} items to the "User wordlist" as a workaround.
P.S. Sorry for sounding quite difficult or hard to understand, I'm dutch by nature.
Betsy25
23rd January 2014, 22:29
One small misbehavior, regarding the "Break long lines" feature, it might be fine to prevent not breaking inbetween a <br /> and a - sign.
For instance, sometimes it tries to break lines like :
Heeft hij onlangs 166 niet gerepareerd?
- Jack heeft onlangs alle drones gerepareerd.
into...
Heeft hij onlangs 166 niet gerepareerd? -
Jack heeft onlangs alle drones gerepareerd.
Nikse555
24th January 2014, 08:53
@Betsy25: I've tried to improve the 'Forrest' issue + tried to fix the 'break' issue too: http://www.nikse.dk/SubtitleEdit.zip
Let me know how it works!
kalehrl
25th January 2014, 20:53
Hi Nikse
I opened a h264 .ts file with a subtitle captured from a satellite and tried to OCR the subtitle. It went fine as far as character recognition is concerned. However, the resulting .srt file had incorrect start time of the subtitle. They were something like 10 hours shifted - first subtitle was 10:00:03,000 instead of 00:00:03,000 and so on. Please try the attached sample to replicate the issue. Subtitle language is Serbian.
https://mega.co.nz/#!l9gCgDDL!vjLjmKS_wnwSyW90EOBRiDF4sHqigWyn-UnnRASdgf8
Also, with the latest version some .mkv files have no sound but I know directshow works fine because mpc plays it.
Nikse555
26th January 2014, 15:52
@kalehrl: thx for the file :)
SE just takes the raw time codes - perhaps subtracting first video time code will give a better time code? Like this: http://www.nikse.dk/SubtitleEdit.zip
About the .mkv files - do you have latest version of LAV filters? https://code.google.com/p/lavfilters/
kalehrl
26th January 2014, 19:42
No, thank you for fixing this issue :)
I use the latest CCCP codec pack which includes LAV Filters 0.60.1.0-22da8ba and MPC-HC 1.7.1.322 (shows up as 333):
http://www.cccp-project.net/forums/index.php?topic=7105.0
jinkazuya
27th January 2014, 19:21
possible to add Chinese & Japanese spellchecking and ORC auto correction please?
Also it is possible to add words to the dictionaries so in the future we don't have to retype or fix those spellchecking or word errors?
Betsy25
30th January 2014, 18:41
@Betsy25: I've tried to improve the 'Forrest' issue + tried to fix the 'break' issue too: http://www.nikse.dk/SubtitleEdit.zip
Let me know how it works!
Sorry for the late reply.
'Forrest' replace issue seems to work fine :o
Regarding the break issue, it doesn't show up in the 'Fixes' list when doing Fix common errors, however it doesn't even show up when i have the max. line length as low as for example 10 in the settings page. It's still flagged in red in the regular window though, but it doesn't appear in the "Fixes" window. Is this how it supposed to be ?
EDIT: This was the "problem" line :
Heeft hij onlangs 166 niet gerepareerd?
- Jack heeft onlangs alle drones gerepareerd.
BTW : The Dutch language uses a lot of dash-connected words, like "mede-speler" (fellow player), "gevechts-troepen" (fighting troops) etc...etc... Perhaps the spell checker can take that into consideration when the dictionary is set to Dutch ? (Right now, we get popups about what to do with "gevechts" but we need to be able to either replace- , add to user dictionary, or check the Dutch dictionary for, the whole dash-connected word.
( EDIT: I posted this on the SE Issue Tracker : issue 207 )
P.S. : Would it be a good idea to have a direct icon in the default window for the "Fix Common Errors..." function ? I see someone had already asked this, in the issue 122.
Nikse555
4th February 2014, 18:44
@kalehrl: I just use quarts.dll for video player... and for me lav-fitlers works fine alone.
@jinkazuya: Is a hunspell dictionary available? Or some open source code (with c# bindings)?
@Betsy25: SE spell check already tries to spell check dash-connected words.
I'll add a "Fix common errors" icon to the main window if someone can make (or find) an icon that matches with the other icons...
Also, SE has moved to GitHub: https://github.com/SubtitleEdit/subtitleedit
(as google has dis-continued downloads on code.google.com)
mood
4th February 2014, 23:15
icons for commonly used tasks in main window
Fix Common Errors, Ms Word spell check and for window Split Long Lines
will be great if you add this icons to main window
Nikse555
5th February 2014, 22:24
Latest beta is here (with "Fix common errors" available as toolbar button9: http://www.nikse.dk/SubtitleEdit.zip (SE 3.3.13 should be out soon)
I need icons for the toolbar... I'm just not god with gfx ;)
lansing
7th February 2014, 20:16
I've used the Microsoft Office Document Imaging for OCR on chinese traditional character with another subtitle software called IdxSubOcr, and the accuracy is close to 99%. However when I use the same option in subtitle edit, the accuracy is less than 5%, most of the time it didn't regconize the characters, and the process is VERY slow. I tried changing the image palette, but there's hardly any improvement.
Nikse555
7th February 2014, 22:21
@lansing: I don't read chinese, so perhaps you can find and compare the source code with SE? Perhaps they do some image scaling or some other tricks?
I did not have MODI installed, but I do now - http://www.microsoft.com/en-us/download/details.aspx?id=21581 (customize setup and choose only modi)
If someone want to have a go at improving this, you can create a fork of the SE source code on GitHub: https://github.com/SubtitleEdit/subtitleedit/fork
lansing
10th February 2014, 17:14
unfortunately I'm not a coder and IdxSubOcr doesn't seem to be open source.
johner23
12th February 2014, 18:15
Hello, dear all.
@lansing: does IdxSubOcr has an english version? Or an english translation? I could find it, but it's in chinese language. And I can't undestand chinese.
Thanks.
DMD
23rd February 2014, 21:47
SORRY
Edit...
Ghitulescu
9th March 2014, 20:13
The above version should be able to open and extract bitmap based subtitles from .ts files... and it might work with .m2ts files as well (just use File -> Open).
Please test :)
I did this today. However I couldn't find any option nor solution on how to save them as such and not OCRed.
Nikse555
9th March 2014, 21:21
I did this today. However I couldn't find any option nor solution on how to save them as such and not OCRed.
Yes, this is not very obvious... but try to right-click in the list view (also check the attached screenshot).
So the subtitle import from the .ts file worked? :)
Ghitulescu
10th March 2014, 08:08
Yes, this is not very obvious... but try to right-click in the list view (also check the attached screenshot).
So the subtitle import from the .ts file worked? :)
Yes, it recognised the bitmaps :) but I did not scroll enough to see whether the colours have been kept (ARTE HD uses various colours for different characterrs so a blind could see who's talking).
I'll wait for the image ... it's not yet approuved.
Ghitulescu
12th March 2014, 17:43
Well, this is what I did (before the image has been approved). And it worked.
It remains to see what all these options do (like Transparent background) ie they affect the display only or they are saved withing the subtitles.
Anyway a big thank, the SUP saved by your software was recognised by many tools (like tsmuxer or bdsup2sub) unlike the one saved by ProjectX.
Music Fan
14th March 2014, 19:43
I was going to ask how to open DVB-SUB included in TS but I just found, I post it for those who search : file, open, file type, all files, choose ts, open.
Ghitulescu
17th March 2014, 10:12
I think that drag'n'drop works too.
Music Fan
17th March 2014, 18:18
Right, I didn't think to this ;)
Music Fan
17th March 2014, 19:21
I discovered a little OCR bug in french : the t is sometimes (at least one time) seen as a l while it is well detected as a t in another line of the same subtitle (same font, same size, exactly the same t). It was in the word "tes" (which mean yours).
@ Nikse555 : it's in the ts file I sent you a few weeks ago, you can test it if you still have it.
von Suppé
17th March 2014, 19:52
... the t is sometimes (at least one time) seen as a l while it is well detected as a t in another line of the same subtitle (same font, same size, exactly the same t).
Can it be that different adjacent characters/letters/signs also influence OCR sometimes? I tend to feel that way.
Music Fan
17th March 2014, 20:19
Probably ;
défaite : ok
tes : ko (seen as "les")
Astonishing because in both cases, the t is followed by a e.
I don't know if a dictionary is used during OCR, if yes I understand that défaite is not seen as défaile because défaile does not exist in french, while les and tes both exist.
von Suppé
29th March 2014, 09:27
Nikse555, is it much work to make SE be able to remember, and preferably save and load settings in the SUP export window? It doesn't seem to remember all the settings, even within the same session. At least framerate and shadow alpha channel are defaulted back every time.
minhjirachi
31st March 2014, 07:44
The function: "Export BDN XML/PNG" doesn't save the setting. So please fix it's problem.
Thank you so much.
Nikse555
13th April 2014, 13:46
Sorry, I cannot change a lot about how Tesseract works - for more info about Tesseract go here: https://code.google.com/p/tesseract-ocr/
I'll try to make SE remember all values from export in next version...
SE 3.3.15 is out - sub/idx files created by SE should now work in handbrake + gpac/mp4box - thx Ryan for fixing this (added an extra byte of value 255 in image data :)
von Suppé
14th April 2014, 08:01
Thanks Nikse555, much appreciated :)
Ghitulescu
14th April 2014, 08:29
To make it as close to perfection as it may be, it may also remember the position of the DVB subtitles (by default in the middle). Sometimes the subtitles are placed under the character that speaks, to help better identifying it. Others use different colours.
Thanks :)
dsmbr
19th April 2014, 00:12
Windows 8.1 (x64) - German
Subtitle Edit 3.3.15 (NET4) --->
1) NET 2-3.5 crashed on starting OCR via Tesseract, I had to switch to the NET4-version. Doesn't make any sense since all .NET-versions are included in Windows 8.1
2) Error on using OCR via image compare.
System.IO.DirectoryNotFoundException: Ein Teil des Pfades "C:\Users\myusername\AppData\Roaming\Subtitle Edit\Ocr\German_Images.db" konnte nicht gefunden werden.
bei System.IO.__Error.WinIOError(Int32 errorCode, String maybeFullPath)
bei System.IO.FileStream.Init(String path, FileMode mode, FileAccess access, Int32 rights, Boolean useRights, FileShare share, Int32 bufferSize, FileOptions options, SECURITY_ATTRIBUTES secAttrs, String msgPath, Boolean bFromProxy)
bei System.IO.FileStream..ctor(String path, FileMode mode, FileAccess access, FileShare share, Int32 bufferSize, FileOptions options, String msgPath, Boolean bFromProxy)
bei System.IO.FileStream..ctor(String path, FileMode mode)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.SaveCompareItem(NikseBitmap newTarget, String text, Boolean isItalic, Int32 expandCount)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.SplitAndOcrBitmapNormal(Bitmap bitmap, Int32 listViewIndex)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.MainLoop(Int32 max, Int32 i)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.mainOcrTimer_Tick(Object sender, EventArgs e)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.ButtonStartOcrClick(Object sender, EventArgs e)
bei System.Windows.Forms.Control.OnClick(EventArgs e)
bei System.Windows.Forms.Button.OnClick(EventArgs e)
bei System.Windows.Forms.Button.OnMouseUp(MouseEventArgs mevent)
bei System.Windows.Forms.Control.WmMouseUp(Message& m, MouseButtons button, Int32 clicks)
bei System.Windows.Forms.Control.WndProc(Message& m)
bei System.Windows.Forms.ButtonBase.WndProc(Message& m)
bei System.Windows.Forms.Button.WndProc(Message& m)
bei System.Windows.Forms.Control.ControlNativeWindow.OnMessage(Message& m)
bei System.Windows.Forms.Control.ControlNativeWindow.WndProc(Message& m)
bei System.Windows.Forms.NativeWindow.Callback(IntPtr hWnd, Int32 msg, IntPtr wparam, IntPtr lparam)
I had to create this missing file manually.
3) Installed German dictorary, started OCR via Tesseract -> error
System.NullReferenceException: Der Objektverweis wurde nicht auf eine Objektinstanz festgelegt.
bei Nikse.SubtitleEdit.Forms.VobSubOcr.OcrViaTesseract(Bitmap bitmap, Int32 index)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.MainLoop(Int32 max, Int32 i)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.mainOcrTimer_Tick(Object sender, EventArgs e)
bei Nikse.SubtitleEdit.Forms.VobSubOcr.ButtonStartOcrClick(Object sender, EventArgs e)
bei System.Windows.Forms.Control.OnClick(EventArgs e)
bei System.Windows.Forms.Button.OnClick(EventArgs e)
bei System.Windows.Forms.Button.OnMouseUp(MouseEventArgs mevent)
bei System.Windows.Forms.Control.WmMouseUp(Message& m, MouseButtons button, Int32 clicks)
bei System.Windows.Forms.Control.WndProc(Message& m)
bei System.Windows.Forms.ButtonBase.WndProc(Message& m)
bei System.Windows.Forms.Button.WndProc(Message& m)
bei System.Windows.Forms.Control.ControlNativeWindow.OnMessage(Message& m)
bei System.Windows.Forms.Control.ControlNativeWindow.WndProc(Message& m)
bei System.Windows.Forms.NativeWindow.Callback(IntPtr hWnd, Int32 msg, IntPtr wparam, IntPtr lparam)
Nikse555
29th April 2014, 17:59
Hi guys,
The "OCR via image compare" is surely broken in SE 3.3.15... sorry about that. I was working on better image->letter splitter + more compact/faster file format.
The Tesseract should still work though... just don't install Tesseract via the installer: http://www.nikse.dk/SubtitleEdit/Help#issues ;)
Tesseract should not have anything to do with the .net framework - and win 8 strangely only included .net framework 4.5 last I checked.
(programs compiled to .net framework 2.0 works on machines with .net framework 2-3.5, programs compiled with .net 4 works on machines with .net framework 4-45 - but perhaps .net programs can compile to native soon: http://msdn.microsoft.com/en-US/vstudio/dn642499.aspx )
von Suppé
30th April 2014, 07:00
Hi Nikse555
I encountered playback issues in MPC-HC with SUP files in mkv container. Within the same sup, some lines are displayed, others not. Now, discarding if MPC-HC has some subtitle issues or not, doing some testing I did find that the SUP output of SE is not consistent.
I exported an srt file as SUP several times - with exactly the same export settings of course. The output files show different hash-check numbers. Now, I do not know if this has something to do with the problems in MPC-HC.
I checked several same outputs of EasySUP and GoSUP and all SUP files show the same hash-check numbers.
VLC and my Dune mediaplayer play SE created SUPs without problems though. I also used them while authoring blu-ray. The burned disks play fine; no subtitle problems.
Any thoughts, please?
Thanks in advance :)
Betsy25
13th May 2014, 14:07
Hi Nikse,
- perhaps this asks for a lot of code, but is some kind of "export/import settings" option on the agenda ?
- Another problem, using Dutch subtitles, the "Fix common errors..." (Fix common OCR errors) capitalizes instances where a line starts with 't (dutch abbrev. of the English word It), Example :
't Is geen rugzaktoerist, hè?
't Was alsof ik voor 'n afgrond stond.
get replaced by...
'T Is geen rugzaktoerist, hè?
'T Was alsof ik voor 'n afgrond stond.
(when dutch language, 't abbrevs always are lowercase. When happening at the start of a sentence, the next following word must be capitalized)
:o
Ghitulescu
13th May 2014, 15:16
I remember a strange behaviour of SE 3.3.13
when loading a .M2TS file with another ending, as it comes from the receiver, it yields an error message that the file is too big.
however, should I replace the ending of thet file with .TS it works as intended.
Betsy25
16th May 2014, 13:25
Doh, Sorry for the incompleteness Nikse,
For dutch language, the same (as 2 posts above) is actually true for the following conditions (same rules apply as for 't abbrevs at start of sentence)
'k : (abbrev. for "Ik" (in English : "I" ) Example : 'k Heb = I have
'm : (special abbrev. for "Hij" (in English : "He" ) Example : 'm Heeft = He has
'n : (abbrev. for "Een" (in English : "An" ) Example : 'n Appel = An apple
'r : abbrev. for "Haar" (in English : "Her" ) Example : 'r Haar = Her hair
That's all, no other special abbrevs. exist in dutch language ;)
jmartinr
17th May 2014, 20:41
Doh, Sorry for the incompleteness Nikse,
For dutch language, the same (as 2 posts above) is actually true for the following conditions (same rules apply as for 't abbrevs at start of sentence)
'k : (abbrev. for "Ik" (in English : "I" ) Example : 'k Heb = I have
'm : (special abbrev. for "Hij" (in English : "He" ) Example : 'm Heeft = He has
'n : (abbrev. for "Een" (in English : "An" ) Example : 'n Appel = An apple
'r : abbrev. for "Haar" (in English : "Her" ) Example : 'r Haar = Her hair
That's all, no other special abbrevs. exist in dutch language ;)
And of course:
's Avonds
Betsy25
17th May 2014, 22:37
And of course:
's Avonds
Yep, but Nikolaj already took care (https://github.com/SubtitleEdit/subtitleedit/commit/1f0eddc66ff8697a972e24dfeb08914609c36ea9) of it also ! :)
Latest version of SE (installer version only) uses the roaming profile.
I only rarely update my installation (why update when it works so well), so I just recently experienced this. My installation, which is NOT from the installer, but copied from the portable zip, also now uses roaming profiles. Any possibility of adding the option of a truly portable setup again, using only files/folders in the installation directory?
I like having parallel installations, in case a new install breaks something, and after installing 3.3.15 (from zip), I cannot start my 3.3.14 installation (image OCR is broken in 3.3.15). 3.3.14 barfs on the 3.3.15 config file. See attachment.
@Nikse555
what differentiates choose "Box for each line" or "one box" on border style?
I see no distinction when export, I have black box in all xml/png regardless of which option I choose.
It was nice to see that too in vobsub export option
But I think it is better to add this option as an option in the context menu as a tag like italic or bold tag and then when export read this tag and apply the style like italic or bold do when export.
Nikse555
27th May 2014, 18:05
@jsa: SE portable should not use the roaming profile... unless you install it under "Program files".
@mood: "Box for each line" makes a box behind each lines where "One box" just makes one box - you can only tell the difference if you have multiple lines.
OK, I've added "One box" for vobsub too.
It's always been possible to make boxing via SSA/ASS for single lines, but I've added it in the context menu for the list view too (right click in the list view in the export window).
thx for testing :)
Compile it from Github or try latest beta here: http://www.nikse.dk/SubtitleEdit.zip
@jsa: SE portable should not use the roaming profile... unless you install it under "Program files".
@mood: "Box for each line" makes a box behind each lines where "One box" just makes one box - you can only tell the difference if you have multiple lines.
OK, I've added "One box" for vobsub too.
It's always been possible to make boxing via SSA/ASS for single lines, but I've added it in the context menu for the list view too (right click in the list view in the export window).
thx for testing :)
Compile it from Github or try latest beta here: http://www.nikse.dk/SubtitleEdit.zip
I have tested but when export not export the box.
If you select other line in list view the "one box" disappears even still the tag <BoxSingleLine> and not see the box in exported sub.
when I say add this to context menu a mean in main list view and not in list view on export window.
adding this to the list view of export window it is not practical
@jsa: SE portable should not use the roaming profile... unless you install it under "Program files".
I originally copied the file structure from the zip file to:
C:\Program Files (x86)\Utils\SubtitleEdit\3.3.15\
This is under 'program files', but not where a normal install would place it. This used to work with local settings, but does not any more.
I tried moving the folder to a different place (not under 'program files'), and now a local settings.xml file is used. :thanks:
minhjirachi
5th June 2014, 09:31
I have a little problem when export the subtitle. When I watch the movie (Avatar) with .srt subtitle file, it match at all period. But when I export the subtitle to .sup or xml/png to muxing in Scenarist, the subtitle not match anymore. At the beginning of the movie, it's match. At the middle, the subtitle was late about 1s. And at the end of the movie, the subtitle was late about 3s or 4s.
So please fix that problem.
P/S: the subtitle edit still not save the configuration in export xml/png mode. I will try the latest version and report later.
von Suppé
6th June 2014, 06:43
Hi Nikse555
Did you read post# 149?
von Suppé
minhjirachi
6th June 2014, 09:21
No more problem about time drift on the latest beta version.
Nikse555
12th June 2014, 18:47
@von Suppé: I think I found the bug - a threading issue that could cause two different bitmaps to write to same header... which is of course very very bad!
C# source code change is here: https://github.com/SubtitleEdit/subtitleedit/commit/61ebdd92ea28cb959d1fc8582e5c23c1fc09b50e
Test version here: http://www.nikse.dk/SubtitleEdit.zip (.net 4, current source, portable version)
@mood: You can use ssa/ass for boxes in the main window
von Suppé
13th June 2014, 10:35
:thanks: Nikse555. Will try test version.
Music Fan
13th June 2014, 15:03
Hi Nikse555,
I tried to open a Hd-dvd sup file but it didn't work ; is this format supposed to be supported by SE ?
mariner
13th June 2014, 16:40
Greetings.
The subtitle duration change tool seems to limit the amount to 1sec. Is there a way to get around this?
If not, would anyone be kind enough suggest alternatives? Something simple for srt subs would suffice.
Thank you and best regards.
Nikse555
13th June 2014, 17:33
@Music Fan: Sorry, I never got around to it...
@marianer: I've added a few more values: http://www.nikse.dk/SubtitleEdit.zip (portable version, beta), or you could use percent or re-calculate durations using chars per sec
mariner
14th June 2014, 14:53
@marianer: I've added a few more values: http://www.nikse.dk/SubtitleEdit.zip (portable version, beta), or you could use percent or re-calculate durations using chars per sec
Thanks Nik.
It would be great if you could also include +10.
Many thanks for providing this extremely useful tool.
Music Fan
14th June 2014, 15:24
@Music Fan: Sorry, I never got around to it...
Do you believe you could add this format or you won't ?
Betsy25
16th June 2014, 18:51
The latest testversion from http://www.nikse.dk/SubtitleEdit.zip gives the following error when being run on Windows 7 32bit.
--> Screenshot (http://i.imgur.com/t2eHAQZ.png)
I forgot to make a backup of the previous testversion, anyone know of a working recent testversion please ?
Nikse555
16th June 2014, 20:01
@mariner: np, added here: http://www.nikse.dk/SubtitleEdit.zip
@Music Fan: I don't have any sample files - but it's probably not something that's very easy to add, but I don't mind taking a look if you think it will be useful (this format is probably dead, right?)
@Betsy25: Ups yes, try the above version.
mariner
18th June 2014, 13:23
@mariner: np, added here: http://www.nikse.dk/SubtitleEdit.zip
.
Thanks Nik. +10 works fine.
Having a little trouble with BD sup creation. It seems PotPlayer/PDVD only displays the sup created using Simple Rendering, while MPC-BE doesn't not have such issue. Appreciate if you could look into it.
Finally, how does one play with the Alpha setting?
Many thanks and best regards.
Nikse555
23rd June 2014, 20:23
Having a little trouble with BD sup creation. It seems PotPlayer/PDVD only displays the sup created using Simple Rendering, while MPC-BE doesn't not have such issue. Appreciate if you could look into it.
Finally, how does one play with the Alpha setting?
PotPlayer works with 'normal' rendering too - if alpha is 255 (fully visible).
You can now change alpha in the color picker dialog :)
mariner
24th June 2014, 06:17
PotPlayer works with 'normal' rendering too - if alpha is 255 (fully visible).
You can now change alpha in the color picker dialog :)
Thanks for the reply, Nik.
1. It appears "normally rendered" sups only work with PotPlayer if Shadow Width is set to 0. Alpha setting doesn't seem to have any effect. Another interesting observation: if fed into BDSup2Sub and saved without any changes, the new sup works.
2. The 3.3.15 seems to have problem retaining "960x540" video resolution setting. The XML output shows "640x272".
3. Is there a way for SubEdit to transform the following timecoed format into proper srt sub with a fixed duration?
00:00:02.000
aaaa
bbbb
.
.
Many thanks and best regards.
Nikse555
24th June 2014, 07:11
1) I'm pretty sure it's an alpha setting thing... and I'm also pretty sure that next version of PotPlayer will be able to play bdsup files with partly transparent shadow or box :)
2) What and how exactly?
3) Post or email a more complete file and I'll take a look.
Ghitulescu
24th June 2014, 07:51
Just a side remark - test the software against hardware players, otherwise you may end with a product that is custom-taylored to a player (in this case PotPlayer) that tomorrow may change.
mariner
24th June 2014, 12:07
1) I'm pretty sure it's an alpha setting thing... and I'm also pretty sure that next version of PotPlayer will be able to play bdsup files with partly transparent shadow or box :)
Attached are three sup files:
0.sup has 0 shadow width and works with PorPlayer,
1.sup has SW=1 Alpha=255, and does not work,
1_exp.sup is processed by BDSup2Sub and works.
14253
2) What and how exactly?
This is the setting for 960x540, and the XML shows 640x272.
14252
<?xml version="1.0" encoding="UTF-8"?>
<BDN Version="0.93" xmlns:xsi="http://www.w3.org/2001/XMLSchema-instance" xsi:noNamespaceSchemaLocation="BD-03-006-0093b BDN File Format.xsd">
<Description><Name Title="subtitle_exp" Content="" /><Language Code="eng" />
<Format VideoFormat="640x272" FrameRate="25" DropFrame="False" />
<Events Type="Graphic" FirstEventInTC="00:00:02:00" LastEventOutTC="00:10:01:13" NumberofEvents="4" />
</Description><Events><Event InTC="00:00:02:00" OutTC="00:00:12:00" Forced="False">
<Graphic Width="470" Height="260" X="245" Y="260">0001.png</Graphic>
</Event>
<Event InTC="00:04:41:12" OutTC="00:04:51:12" Forced="False">
<Graphic Width="470" Height="194" X="245" Y="326">0002.png</Graphic>
</Event>
<Event InTC="00:05:55:00" OutTC="00:06:05:00" Forced="False">
<Graphic Width="438" Height="194" X="261" Y="326">0003.png</Graphic>
</Event>
<Event InTC="00:09:51:13" OutTC="00:10:01:13" Forced="False">
<Graphic Width="437" Height="194" X="261" Y="326">0004.png</Graphic>
</Event>
</Events></BDN>
3) Post or email a more complete file and I'll take a look.
Something like this, say for 10s duration.
00:00:02.000
Junior Semifinal, part 1
Aidiba Talamunuer, Berezan
Bogdan Voloshin, Yaroslavl
Alexandr Doronin, Almaty
00:04:41.480
G. Zhubanova
«Kui»
Aidiba Talamunuer, Berezan
00:05:55.000
N. Mendigaliev
«Steppe»
Bogdan Voloshin, Yaroslavl
00:09:51.520
A. Lokshin
«Dance»
Alexandr Doronin, Almaty
Many thanks and best regards.
mariner
24th June 2014, 12:14
Just a side remark - test the software against hardware players, otherwise you may end with a product that is custom-taylored to a player (in this case PotPlayer) that tomorrow may change.
Have not tested on a player yet.
SE has the unique ability to create sups with non standard video resolution. Want to make sure everything is in order.
Music Fan
24th June 2014, 12:26
@Music Fan: I don't have any sample files - but it's probably not something that's very easy to add, but I don't mind taking a look if you think it will be useful (this format is probably dead, right?)
I found another program (BDSup2Sub) able to convert Hd-dvd sup into BD Sup, so if it's not easy to add in SE, don't waste your time.;)
Nikse555
25th June 2014, 19:16
@mariner: Could you test this version: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version) - does the file format work + resolution? (if not, how to you get to the export bdn xml window?)
Notice that both "border" and "shadow" can have alpha adjusted.
@Music Fan: OK, I'll use my time for preparing SE 3.4 instead :)
mariner
26th June 2014, 07:14
@mariner: Could you test this version: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version) - does the file format work + resolution? (if not, how to you get to the export bdn xml window?)
Notice that both "border" and "shadow" can have alpha adjusted.
Thanks for posting build 435, Nik.
1. PotPlayer issue: no change.
I've posted what BDSup2sub reports when reading a sup file. It seems the difference between the bad and good sup file is the update entries. All the good ones, including the one fixed by BDSup2sub, have less than 256 entries. Is this related to the issue at hand?
Bad sup
BDSup2Sub++ 1.0.2 - a converter from Blu-Ray/HD-DVD SUP to DVD SUB/IDX and more
0xdeadbeef, mjuhasz, Adam T.
Official thread at Doom9: https://forum.doom9.org/showthread.php?t=167051
Loading G:/sup/1/j.200.sup
PCS ofs:0x00000000, START, size:0x0013, comp#: 0, Object Id: 0, forced: false, palID: 0, objID: 0
PTS start: 00:00:01.955, screen size: 960*540
WDS ofs:0x00000020, size:0x000a, windows 0 dim: 335*155
PDS ofs:0x00000037, size:0x0502, ID: 0, update: 0, 256 entries
ODS ofs:0x00000546, size:0x5561, img size: 335*155
END ofs: 0x00005ab4
#< 1 (00:00:01.955)
PCS ofs:0x00005ac1, START, size:0x000b, comp#: 1
PTS start: 00:00:11.955, screen size: 960*540
WDS ofs:0x00005ad9, size:0x000a, windows 0 dim: 335*155
END ofs: 0x00005af0
PCS ofs:0x00005afd, START, size:0x0013, comp#: 0, Object Id: 0, forced: false, palID: 0, objID: 0
PTS start: 00:00:21.435, screen size: 960*540
WDS ofs:0x00005b1d, size:0x000a, windows 0 dim: 335*119
PDS ofs:0x00005b34, size:0x0502, ID: 0, update: 0, 256 entries
ODS ofs:0x00006043, size:0x270a, img size: 335*119
END ofs: 0x0000875a
#< 2 (00:00:21.435)
PCS ofs:0x00008767, START, size:0x000b, comp#: 1
PTS start: 00:00:31.435, screen size: 960*540
WDS ofs:0x0000877f, size:0x000a, windows 0 dim: 335*119
END ofs: 0x00008796
"Fixed" by BDSup2sub
BDSup2Sub++ 1.0.2 - a converter from Blu-Ray/HD-DVD SUP to DVD SUB/IDX and more
0xdeadbeef, mjuhasz, Adam T.
Official thread at Doom9: https://forum.doom9.org/showthread.php?t=167051
Loading G:/sup/1/j.200_exp.sup
PCS ofs:0x00000000, START, size:0x0013, comp#: 0, Object Id: 0, forced: false, palID: 0, objID: 0
PTS start: 00:00:01.960, screen size: 960*540
WDS ofs:0x00000020, size:0x000a, windows 0 dim: 336*156
PDS ofs:0x00000037, size:0x04f8, ID: 0, update: 0, 254 entries
ODS ofs:0x0000053c, size:0x59ea, img size: 336*156
END ofs: 0x00005f33
#< 1 (00:00:01.960)
PCS ofs:0x00005f40, START, size:0x000b, comp#: 1
PTS start: 00:00:11.960, screen size: 960*540
WDS ofs:0x00005f58, size:0x000a, windows 0 dim: 336*156
END ofs: 0x00005f6f
PCS ofs:0x00005f7c, START, size:0x0013, comp#: 2, Object Id: 0, forced: false, palID: 0, objID: 0
PTS start: 00:00:21.440, screen size: 960*540
WDS ofs:0x00005f9c, size:0x000a, windows 0 dim: 336*120
PDS ofs:0x00005fb3, size:0x04fd, ID: 0, update: 0, 255 entries
ODS ofs:0x000064bd, size:0x296b, img size: 336*120
END ofs: 0x00008e35
#< 2 (00:00:21.440)
PCS ofs:0x00008e42, START, size:0x000b, comp#: 3
PTS start: 00:00:31.440, screen size: 960*540
WDS ofs:0x00008e5a, size:0x000a, windows 0 dim: 336*120
END ofs: 0x00008e71
2. XML video format issue:
The 960x540 format is now correctly shown, but 1280x720 is shown as 1080p.
3. XML timecode for 23.976, 29.97 and 59.94 fps.
I notice SE does not follow the usual convention of speeding up the XML timecode by a factor of 1.001 when the fps is one of the above mentioned. So when the XML produced by SE is used by BDSup2sub, the resulting BD sup is slowed down by a factor of 1.001.
Any good reason for doing this?
Many thanks and best regards.
Nikse555
26th June 2014, 14:28
1) Thx a lot for the info - yes, one less entry in the color palette seems to fix this :)
http://www.nikse.dk/SubtitleEdit.zip (new beta, portable version, )
2 would you rather have just the resolution? (I think I saw 1080p in a bd xml file)
3 Well, I mostly work with srt... but I've given it a try for File -> Export -> BDN xmp/png (does it work in above beta version?)
mariner
27th June 2014, 09:35
1) Thx a lot for the info - yes, one less entry in the color palette seems to fix this :)
http://www.nikse.dk/SubtitleEdit.zip (new beta, portable version, )
255 appears to do the trick. Thanks.
What's wrong with 256?
2 would you rather have just the resolution? (I think I saw 1080p in a bd xml file)
I would think you want the subtitle to follow the resolution of the video?
3 Well, I mostly work with srt... but I've given it a try for File -> Export -> BDN xmp/png (does it work in above beta version?)
No change.
4. Timecode import.
Thanks for adding the feature. Is there a way to make the duration fixed?
Many thanks and best regards.
Nikse555
30th June 2014, 19:13
I've no clue what's wrong with 256... it's the max number of values for a byte so 256 sounds more correct than 255 - but I've never seen the bluray sup specs :(
2) Thx, my bad (reading)... should be fixed in latest beta - http://www.nikse.dk/SubtitleEdit.zip
3) If you choose those drop-frame time codes the times should now be multiplied by 1.001... I think.
4) One way could be Tools -> Adjust durations and add 10 seconds, the use Tools -> Apply duration limits (with 10 seconds as max duration)
Superb
30th June 2014, 22:50
Without reading the rest of the issue (I've only read the first line of the comment above me):
255 is the maximum number for a byte. Don't forget that the counting starts from 0.
So 0-255 are 256 different values (or 2^8; the number of unique binary combinations of 8 bits).
Nikse555
1st July 2014, 16:11
@Superb: Yes, I was unclear... now SE creates index 0-254 palettes for bluray sup (255 values), and ealier SE created 0-255 index palettes (256 different palettes), but that caused some programs not to display the subtitles... I would like to know what the specification says ;)
mariner
2nd July 2014, 06:44
I've no clue what's wrong with 256... it's the max number of values for a byte so 256 sounds more correct than 255 - but I've never seen the bluray sup specs :(
2) Thx, my bad (reading)... should be fixed in latest beta - http://www.nikse.dk/SubtitleEdit.zip
3) If you choose those drop-frame time codes the times should now be multiplied by 1.001... I think.
Perhaps you'd like to discuss these matters with ps auxw. He'd e also be able to advise you on how to render contiguous subtitles without flickering.
4) One way could be Tools -> Adjust durations and add 10 seconds, the use Tools -> Apply duration limits (with 10 seconds as max duration)
Brilliant.
Many thanks and best regards.
Music Fan
2nd July 2014, 11:16
I noticed a little problem with VLC when playing video including Blu-ray sup : if there is no space (in time) between 2 successive subtitles in the original srt, the sup are not correctly displayed.
Example ;
1
00:01:48,955 --> 00:01:51,082
Hello guys,
2
00:01:51,082 --> 00:01:53,300
how are you ?
I guess the solution is to shorten the first subtitle and get this (for example) ;
1
00:01:48,955 --> 00:01:50,900
Is it possible to detect it with SE (for srt) and shorten the subtitles only when their end timecode is the same than the debut timecode of the following subtitle ?
I could apply duration limits but I'm afraid that some non problematic subtitles would become too short.
minhjirachi
4th July 2014, 11:43
When I export the subtitle as xml/png, I can't import that xml to Scenarist. So please fix that problem.
Thank you so much.
von Suppé
6th July 2014, 23:21
Hi Nikse555
First tests with latest version show that the hash-check numbers of multiple, same exports to SUP format are equal. Thanks for this fix.
Music Fan
14th July 2014, 01:29
Is there a way to convert DVB-SUB (from HDTV TS) to Blu-ray SUP without doing OCR ?
Because the OCR does not work well with one of my recordings while the original sup is correctly displayed by SE, thus a simple conversion is enough if possible.
kalehrl
14th July 2014, 08:11
Is Blu-ray SUP just the ordinary SUP?
If so, then ProjectX converts DVB subtitles to SUP format and also to VobSub.
Music Fan
14th July 2014, 09:03
Thanks, I know, I already tried ProjectX but the conversion is not very well done.
Nikse555
14th July 2014, 09:13
SE 3.4 out :)
@Music Fan: I've added a small overlap fix in 3.4 for outputting bluray sup when end time = next start time, then the end time will be shortened by one ms... but it might not be enough. SE -> Tools -> Min display time between subtitles might help though.
To convert imaged based formats to Blu-ray sup, just right click in the list view in the ocr window and choose Export -> Bluray sup...
@minhjirachi: What's the error message? How can I test it?
@von Suppé: Thx for testing :)
Music Fan
14th July 2014, 10:42
SE 3.4 out :)
@Music Fan: I've added a small overlap fix in 3.4 when outputting bluray sup when end time = next start time, then the end time will be shortened by one ms... but it might not be enough. SE -> Tools -> Min display time between subtitles might help though.
Great, I will have a look on this :)
To convert imaged based formats to Blu-ray sup, just right click in the list view in the ocr window and choose Export -> Bluray sup...
Thanks, very well hidden option ;)
It works but the subtitle size is much smaller than the original (both played by VLC) while the font size can't be changed in SE in this case.
I chose 1080p for resolution because this is the video resolution, but the original sub has maybe a wrong size (for example 576p while it should be 1080p) which could explain the size change when exporting in 1080p.:confused:
How to export in 1080p keeping the same size ?
Nikse555
14th July 2014, 10:46
Sorry, SE don't have any re-sizing of subtitle images, so you would have to do the resizing in BDSup2Sub... or ocr and export.
Music Fan
14th July 2014, 11:17
Ok thanks, and do you know why VLC does not display the original dvb-sub and the blu-ray sup (remuxed with TSMuxer) with the same size ? Could it be related to the SE conversion ?
Music Fan
14th July 2014, 11:54
Other question : is there a way to change sup maximum duration and gap without doing OCR ?
These option are not available in the OCR window, and if I do ok in the OCR window without doing OCR, I get the timecodes without text in the main window, but the maximum duration and gap can be changed there.
Then, if I export in Blu-ray sup, the subtitle are empty. If I save it in srt, I have a file with the corrected timecodes but without text.
If these changes can't be applied in the OCR window, could you add an option to load the timecodes from an external srt (actually the srt without text created just before) and apply it to the sup (if possible) ?
Thanks for your work ;)
Nikse555
14th July 2014, 16:00
I've tried to add an "Import new time codes"... http://www.nikse.dk/SubtitleEdit.zip (portable version)
Does it work?
Music Fan
14th July 2014, 16:50
Yes, you rule !:)
But there is a minor problem : curiously, when I open the sup with modified timecodes (thus after having imported new time codes and exported in Blu-ray sup), the timecodes are all displayed 44 ms sooner than those in the srt used as new time codes.
For example, if the time code of the first line of the good srt is ;
1
00:00:07,921 --> 00:00:12,241
I get this in the corrected sup ;
1
00:00:07,877 --> 00:00:12,197
But as I can change srt's timecode easily with SE to compensate this little change, that's not a big problem.;)
von Suppé
14th July 2014, 17:01
I'm gonna try this out too, sounds promising.
Yes, you rule !:)
But there is a minor problem : curiously, when I open the sup with modified timecodes (thus after having imported new time codes and exported in Blu-ray sup), the timecodes are all displayed 44 ms sooner than those in the srt used as new time codes.
For example, if the time code of the first line of the good srt is ;
1
00:00:07,921 --> 00:00:12,241
I get this in the corrected sup ;
1
00:00:07,877 --> 00:00:12,197
But as I can change srt's timecode easily with SE to compensate this little change, that's not a big problem.;)
Or, after importing the new timecodes, you could also change all their in/outs manually :D
Betsy25
14th July 2014, 20:32
The "Fix common OCR errors" is a bit too greedy, for example in the example below, it will propose to replace "for" with "For", however this should not be the case.
Any idea how to make the method stay away from automatically capitalizing the new subtitle line when the previous one ended with some "end of sentence" string ? There are quite a lot of "false positives" by default.
43
00:08:07,379 --> 00:08:11,033
I warned their president...
- Say it John, Say it !
44
00:08:11,133 --> 00:08:13,173
for doing business with them.
von Suppé
14th July 2014, 22:21
Nikse555, is it possible to add this to the "Fix common error tools":
If there are 2 lines, and both begin with a dash, being able to remove the dash (and the following space, if it's there) in the first line?
This would be as opposite of the tool "Fix (add dash) line pairs with only one dash (-)"
Maybe call it "Fix (remove first dash) line pairs with two dashes" or something?
Example:
- Hi, how are you?
- I'm fine, thanks.
-->
Hi, how are you?
- I'm fine, thanks.
Thanks for your hard work :)
Betsy25
16th July 2014, 20:28
Proposition for the "Fix Unneeded Spaces" :
In case of <somestring><DASH><space><nextstring>, when <nextstring> is "and" or "or", do not remove the <space>.
("en" or "of" in Dutch)
Examples :
(English)
What are your long- and altitude stats ?
There is an X- and a Y-chromosome.
Did you buy that first- or second-handed ?
(Dutch)
Wat zijn je voor- en familienaam ?
Was het in het voor- of najaar ?
I've commited some brewing on my github (https://github.com/Betsy25/subtitleedit/commits/My-fixes2) that probably is far below your coding standards, but perhaps you can refine it.
Ghitulescu
21st July 2014, 07:53
If there are 2 lines, and both begin with a dash, being able to remove the dash (and the following space, if it's there) in the first line?
This would be as opposite of the tool "Fix (add dash) line pairs with only one dash (-)"
Maybe call it "Fix (remove first dash) line pairs with two dashes" or something?
Example:
- Hi, how are you?
- I'm fine, thanks.
-->
Hi, how are you?
- I'm fine, thanks.
)
Why would be this a correct formatting?
Yes, I've seen this also on commercial movies, but I still have no explanation why one line is treated differently than the other one, although both are dialogues ...:confused:
von Suppé
21st July 2014, 20:34
I'm not saying it's correct formatting, if there is any. I and many people just like it this way and find the first line dash not necessary.
I myself do a lot of manual subtitling and there's a lot of feel to it, which I find more important than correct formatting or not. A first dash disturbs my feel to the text in most cases.
Betsy25
21st July 2014, 21:55
I'm not saying it's correct formatting, if there is any. I and many people just like it this way and find the first line dash not necessary.
I myself do a lot of manual subtitling and there's a lot of feel to it, which I find more important than correct formatting or not. A first dash disturbs my feel to the text in most cases.
the - normally "tells" the viewer that a different person than the one who told the last line is now speaking, so it's perfectly normal to have both lines beginning with - in that case. (at least, that's the way I understand it):)
von Suppé
22nd July 2014, 07:54
Yes, I know what you mean.
If one would consist in that way, in every new subtitle display where the speaker of the first line is another than the last one in the previous subtitle display, how many cases would you have to start with a dash to tell the viewer this? It feels unnatural to me to do so.
Example:
Display 1:
Son, as I don't have much time today,
will you do me a favor and wash the car?
Display 2 would have to be like this:
- I will in the afternoon dad, because
I promised mom to fix her computer first.
Or similar things. Doesn't feel right. But everybody should of course do as they see or feel fit.
TheSkiller
22nd July 2014, 13:07
I agree with von Suppé.
Whenever I'm creating subtitles myself I put the dash (which in my humble opinion should always be a so called "em dash" —, not a "hyphen" - ) in front of the second line, in other words exactly where the change of the speaker happens within the same subtitle.
I also think it's easier to follow such a subtitle if both lines are left-aligend but as a whole the subtitle should of course be centered within the video frame, with the top line not starting before the em dash of the bottom line.
http://picload.org/thumbnail/lwgpdww/subtitle_speakerchange_with_em.png (http://picload.org/image/lwgpdww/subtitle_speakerchange_with_em.png)
(click to enlarge)
That's what I personally consider "most aesthetic" and easiest to read.
I'm quite sure the main reason for putting a dash in front of both lines is rather just a technical shortcoming because it is easier to format this way. But of course it also depends on what the viewer is expecting and used to.
mood
23rd July 2014, 08:10
@Nikse
It is possible add boxing options to context menu???
von Suppé
24th July 2014, 09:26
@ mood
Can you tell me what "boxing" is?
mood
26th July 2014, 01:42
@von Suppé
is in version 3.4 change log.
boxing is this:
http://i.imgur.com/xlNf85L.png
you can add tags manually (<BoxSingleLine>) and (<BoxMultiLine> to the lines to get the effect shown in the image above.
is a black box on background subtitle, you can have lines with black box on background and without them in same subtitle.
this only on sub image based.
this is great, but only can do this manually or in export window.
And in export window not work very well because the effect only work in one line whatever more lines you add the effect
I use this in idx/sub and I need this on context menu like italic tag and others
this effect is useful when you have hardsubs on video.
von Suppé
26th July 2014, 11:58
I understand. Thanks for clear explanation. :)
Nikse555
31st July 2014, 12:56
About the dialogues and hyphen, try File - Plugins - "Dialogue one hyphen only".
The rules regarding one or two hyphens might be different depending on language I think...
@mood: In export you can choose default "Border style" incl. "box single line" or "one box".
The list view (and hence the list view context menu) in the export will be improved with Ctrl+a for select all, ctrl+d for de-select + ctrl+shif+i for inverse selection. Changing selection will also be much faster (was very slow previously).
Betsy25
31st July 2014, 16:07
Nikse, you're willing to take a look at my 2 latest brewings here : https://github.com/Betsy25/subtitleedit/commits/My-Fixes-2
I'm a complete noob in C#, but I'm pretty sure you'll understand what I'm trying to do. (Regarding the dash cases followed by and or or)
mood
31st July 2014, 17:28
@Nikse555
thanks, but the "Border style" applies to all lines and I do not want to apply the style to all lines, only the lines that I choose.
In export window the boxing effect only work for one line, if you applie more than one line, the effect you applie on line before disappears, the tags <BoxSingleLine> or <BoxMultiLine> still there but without effect.
Nikse555
31st July 2014, 20:20
@Betsy25: I'll take a look... sorry I've looked earlier (was on vacation).
@mood: thx, a bug - hopefully fixed now (on Github and in this portable beta - http://www.nikse.dk/SubtitleEdit.zip )
mood
31st July 2014, 22:45
@Nikse555
thanks, now work but I hope you add this style to context menu ;)
Betsy25
1st August 2014, 16:04
@Nikse,
when you have the time, will you take a look at your PM's here on Doom9 forum please ?
mariner
3rd August 2014, 16:48
Greetings Nik.
Having trouble getting the BD sup to follow the font settings defined in ass subtitles. The preview panel seems to work fine, but the sup created uses only one font. Is this a known issue?
Many thanks and best regards.
Nikse555
3rd August 2014, 20:27
Having trouble getting the BD sup to follow the font settings defined in ass subtitles. The preview panel seems to work fine, but the sup created uses only one font. Is this a known issue?
Thx for reporting this :)
Should be fixed on github + in this portable beta: http://www.nikse.dk/SubtitleEdit.zip
mariner
4th August 2014, 16:51
1. Thanks for the quick fix, Nik.
2. There appears to be a problem in the handling of overlapping subs, some ending up with 0 duration.
Many thanks and best regards.
Nikse555
4th August 2014, 17:32
1. Thanks for the quick fix, Nik.
2. There appears to be a problem in the handling of overlapping subs, some ending up with 0 duration.
1) np
2) where and doing what?
mariner
4th August 2014, 17:49
It seems the overlapped graphic elements need to be combined into one.
Music Fan
4th August 2014, 19:48
@ Nikse555 : did you find where does this little problem come from ?
there is a minor problem : curiously, when I open the sup with modified timecodes (thus after having imported new time codes and exported in Blu-ray sup), the timecodes are all displayed 44 ms sooner than those in the srt used as new time codes.
For example, if the time code of the first line of the good srt is ;
1
00:00:07,921 --> 00:00:12,241
I get this in the corrected sup ;
1
00:00:07,877 --> 00:00:12,197
But as I can change srt's timecode easily with SE to compensate this little change, that's not a big problem.;)
I talked about this in post #200 ;
http://forum.doom9.org/showthread.php?p=1686798#post1686798
Nikse555
4th August 2014, 20:31
@mariner: Yeah, I guess... perhaps a warning would be nice...
@Music Fan: Hm, perhaps to give the decoder a tiny bit extra time to decode the subtitles? I don't know ;)
The code is here: https://github.com/SubtitleEdit/subtitleedit/blob/master/src/Logic/BluRaySup/BluRaySupPicture.cs#L45
Betsy25
7th August 2014, 22:33
Perhaps a bug.
In case the user has unchecked "Start with uppercase letter after paragraph" in the Fix common Errors dialog (and the other "start with uppercase" choices), some items are still detected and presented to be Uppercased using the "Fix Common Errors (using the OCR replace list)" method.
For example, subtitles 444 and 1047 in the subtitle file below.
444 --> als ie tenminste nog capabel was...
1047 -> b kwadraat?
Subtitle : Dutch .SRT subtitle (https://www.dropbox.com/s/kgocs0sixny85e7/Into%20the%20Wild%20%282007%29%20%281080p%29.ned.srt)
Thunderbolt8
8th August 2014, 14:17
Maybe a bug, for some reason I couldnt delete an entry from the OCR fix list which I had added myself before. had to edit the .xml file to get rid of it.
Nikse555
10th August 2014, 07:12
@Betsy25: Hm, yes. "Fix common OCR errors" also changes first letter to uppercase...
@Thunderbolt8: I could not re-create this... let me know if it occurs again or if you can re-create it.
Also, changes to interjections require the 'Remove text for HI' window to be closed and re-opened for the changes to take effect... will be fixed next update.
Thunderbolt8
11th August 2014, 13:15
in the remove text for hearing impaired menu, would it be possible to add exceptions for removing text between ( ) and { } in case of .ass files? the screen position in some .ass subs is indicated between these markers. perhaps when such marker both occur in the same line or when the ( ) brackets are inside the { } ones?
currently, its like this:
{} ticked: {\an5\pos(958,176)}JACKIE: <i>It seems as though</i> -----> <i>It seems as though</i>
() ticked: {\an5\pos(958,176)}JACKIE: <i>It seems as though</i> -----> <i>It seems as though</i>
--> should be: {\an5\pos(958,176)}<i>It seems as though</i>
{} ticked: {\an4\pos(705,833)}<i>Today really marks</i> -----> <i>Today really marks</i>
() ticked: {\an4\pos(705,833)}<i>Today really marks</i> -----> {\an4\pos}<i>Today really marks</i>
--> should be: {\an4\pos(705,833)}<i>Today really marks</i>
{\an4\pos(556,902)}Huh? -----> {\an4\pos(556,902)}
--> should be: whole line should get deleted because its empt after the position information after SHD removal.
{\an4\pos(814,825)}- [ Panting ] -----> {\an4\pos(814,825)}-
--> should be: whole line should get deleted because theres only a hyphen left after SHD removal, but line doesnt get removed because the subtitle screen position information is considered as text information.
Nikse555
12th August 2014, 16:47
@Thunderbolt: thx for reporting this... I probably use srt files most of the time.
Should be fixed on github + in this portable beta: http://www.nikse.dk/SubtitleEdit.zip (also contains the interjections-fix)
Let me know if you find more stuff :)
Betsy25
16th August 2014, 04:49
@Betsy25: Hm, yes. "Fix common OCR errors" also changes first letter to uppercase...
So, there's no way to stop it from (false positive) uppercasing, without giving up the whole "Fix common Errors" package ?
Thunderbolt8
18th August 2014, 21:06
there are a few things I still noticed. some of them to come at a later point. but for now:
- if there are double hyphens "--" in line initial position already present before SHD removal has been run (e.g. --the wheather will be...) then those double hypens shouldnt be changed and taken into consideration for SHD removal.
e.g. "SPEAKER: --the wheather will be...." in this case the double hyphens should be disregarded from removal. this should only apply with cases in which the double hyphens were present from the start, not as a leftover result from a previous SHD removal (i dont know how the SHD removal stage works, if there are for examples 2 runs you do after another when pressing button only once)
- hyphens and other SHD information in a subtitle line (for .srt) should be taken into consideration if that subtitle line for example consists of two or more lines on screen.
e.g
- It's already done. (- It's already done.|BENJAMIN: You--)
BENJAMIN: You--
gets changed to
- It's already done. (- It's already done.|You--)
You--
while it should get changed to
- It's already done. (- It's already done.|- You--)
- You--
(this example can get fixed though with "fix line beginning with dash (-)" in the fix common errors tab, but that option might not be activated, possibly because it could interfere negatively in other situations within the same file. so it would be nice if this situation could be resolved already during SHD removal)
- same as above in case of .ass subtitles, if you have 2 or more subtitle lines with identical timestamps, meaning they belong together on screen.
e.g.
00:04:40:22 00:04:43:23 {\an4\pos(775,828)}- Hey, hey!
00:04:40:22 00:04:43:23 {\an4\pos(775,900)}- [ Laughing ]
gets changed to
00:04:40:22 00:04:43:23 {\an4\pos(775,828)}- Hey, hey!
while it should get changed to
00:04:40:22 00:04:43:23 {\an4\pos(775,828)}Hey, hey!
- in case of .ass subs apparently brackets like [] or () sometimes arent applied in each line, but each block with same timestamps which belog together
00:29:38:23 00:29:41:11 {\an4\pos(1177,684)}[ Ship's Horn Blows
00:29:38:23 00:29:41:11 {\an4\pos(1178,756)}In Distance ]
this get ignored by SHD removal (and "Fixing missing [ in line" only fixes the 2nd line, but doesnt list the first one for some reason), maybe these situations can be added to SHD removal as well in case of .ass subs.
it might already get tricky now that implementing new rules could interfere with other rules ;)
mariner
20th August 2014, 16:21
Greetings Nik. A few issues with BD sup rendering:
1. In a srt sub, if the italic tag is preceded by space after a line break, the placement of that line is off.
2. Can't seem to get the font formatting defined by AEGISUB style editor to work. Are these supported by SE?
3. Fade doesn't seem to work as well.
Thanks for looking into these issues.
Thunderbolt8
24th August 2014, 20:33
in which file are the user edited/added information of the "Remove interjections" list stored?
Nikse555
24th August 2014, 21:55
in which file are the user edited/added information of the "Remove interjections" list stored?
In Settings.xml... in a tag called "Interjections".
@mariner: thx for the info about (1) italic tag is preceded by space after a line break... will be fixed in next update.
(2) basic stuff like font style, color should work...
(3) I don't know anything about fade.
von Suppé
28th August 2014, 09:26
Hi Nikse555
Is it possible to make SE show the subtitles as SUP preview in the videopreview window? Preferably with being able to adjust SUP settings (font, size, position etc.) for all - or selected only.
I think I'm asking a lot here... :rolleyes:
cheers
von Suppé
Thunderbolt8
28th August 2014, 20:13
the "Fix common errors" window always opens on my primary screen, even though I have the program on my 2nd screen and dragged the window over already before. could you please fix that? other windows also open at that screen at which they have been closed down last.
Thunderbolt8
4th September 2014, 16:56
I added "Ha!" to the remove interjections list, but subtitle lines just consisting of "Ha!" dont get removed by that. bug?
edit: well it works if I dont include the exclamation mark.
edit²: but then again a line consisting of "- Ha!", resulting from a previous remove interjections run, does not get removed.
Music Fan
8th September 2014, 16:07
rev 3.4.2 ;
https://github.com/SubtitleEdit/subtitleedit/releases
kalehrl
12th September 2014, 21:59
Hi Nikse
I keep getting this error:
Message: Exception from HRESULT: 0x80040265
Source: SubtitleEdit
StackTrace:
at QuartzTypeLib.IMediaControl.RenderFile(String strFilename)
at Nikse.SubtitleEdit.Logic.VideoPlayers.QuartsPlayer.Initialize(Control ownerControl, String videoFileName, EventHandler onVideoLoaded, EventHandler onVideoEnded)
at Nikse.SubtitleEdit.Logic.Utilities.InitializeVideoPlayerAndContainer(String fileName, VideoInfo videoInfo, VideoPlayerContainer videoPlayerContainer, EventHandler onVideoLoaded, EventHandler onVideoEnded)
I already have the latest LAV Filters installed.
Nikse555
13th September 2014, 06:25
Hi Nikse
I keep getting this error:
I already have the latest LAV Filters installed.
Is it only one video file that gives this error or all videos files?
LAV Filters cannot decode all video files so you might switch to VLC media player for special files like 10-bit color or similar.
Also, from version 3.4.0 SE runs 64-bit on 64-bit operating systems which will require 64-bit lav filters too.
kalehrl
13th September 2014, 07:27
I have Combined Community Codec Pack installed and I guess it ships 32bit LAV filters.
I installed 64bit version of the filters, and now SE works.
Thanks
Thunderbolt8
15th September 2014, 20:37
theres still problem with hearing impaired stuff removal in .ass subs and the "remove text before a colon" option. when the "only if text is UPPERCASE" box is ticked as well, then the information about the speaker is not removed:
{\an4\pos(1335,891)}NIC: Shh! --> {\an4\pos(1335,891)}NIC:
or {\an4\pos(691,748)}WOMAN: (OVER PA) --> {\an4\pos(691,748)}WOMAN:
(both lines should be removed entirely as they dont include any speech)
and when I leave the "only if text is UPPERCASE" box unticked then everything before the colon is removed including the screen position information part of the line:
{\an4\pos(652,817)}NIC: "Your hack into MIT --> "Your hack into MIT
(screen position information needs to be preserved and not taken into consideration for removal if there is still speech present in the line)
minhjirachi
21st September 2014, 03:31
Does anyone know which software can make the effect for the subtitle? Like Fade In Effect or Fade Out Effect.
Superb
21st September 2014, 05:55
Does anyone know which software can make the effect for the subtitle? Like Fade In Effect or Fade Out Effect.Aegisub... for example... http://docs.aegisub.org/3.2/ASS_Tags/
tuco76
24th September 2014, 13:12
Does anyone know which software can make the effect for the subtitle? Like Fade In Effect or Fade Out Effect.
PotPlayer is sort of disneyland for the habitual subtitle tinkerer ;)
Standard Fade effect, Fade in/ out adjustable strengt 1-100%,
and an almost ridiculous lot of else subtitle settings, selections/ adjustments/features/gimmicks on top of that. Be aware! :)
You prefer it rather basic and readily comprehensible-
better dont bother with :)
Mole
25th September 2014, 13:28
Do you think it'll be possible to open several subs at the same time?
For example, sometimes I work with subs from TV series, and I'd like to be able to spell check them all with exactly the same settings, so that certain weird name not in the dictionary can be skipped automatically on all the subs or certain words be renamed on all subs. But I'd rather not add them all permanently in the program because occasionally there may be spelling variations on a particular name, which may be used in one movie, but spelled differently in another. So I want to manually specify which spelling I want to use.
This is often the case for us who work with English subs for non English movies.
Maybe each sub can be in it's own tab or something.
Also, you should separate word list which we have added manually into a dedicated XML file, so that each time we upgrade your program we can simply just copy this XML file over.
Currently all are stored in the eng_OCRFixReplaceList and each time I upgrade your program, I need to manually add in my own stuff in there, because each new version will write over this file.
Problem is that all the words will be sorted alphabetically in eng_OCRFixReplaceList, and often I don't remember which one was my own.
Hope you understand what I'm trying to say.
minhjirachi
26th September 2014, 18:22
I think the subtitle edit have a little bug with sub color. Check this script:
VAV STUDIO Hân Hạnh Giới thiệu
Bộ phim: <font color=yellow>"CUỘC CHIẾN LUÂN HỒI"</font>
Export as XML/PNG and *.SUP will have different results. XML/PNG get the correct color but can't import to Scenarist without BDSup2Sub and have a little bug about time stamp. Export as *.SUP will have bug about the sub color.
Hope you can fix it as soon as possible.
Music Fan
3rd October 2014, 13:49
Hi Nikse,
I have a lot of problems with DVB-SUB (included in TS) ;
1)a lot of ts can't be opened ("unhandled exception")
2)with those which are opened, the size after conversion in Blu-ray sup without OCR is completely different (much more little) than what I see when I play the original TS in MPC-HC and VLC (which both handle DVB-SUB). But for SE the size of the lines are the same before and after conversion : I mean that just after parsing the ts, if the detected height of a line is 33 (in the OCR window), the size will still be 33 for this line when opening the Blu-ray sup created from this ts, thus I don't know what is the real size and why they are different for MPC-HC and VLC while they are the same for SE.:confused:
And when put on AVCHD and played with a standalone player, the size of the Blu-ray sup is also very little, same size than what I see in MPC-HC and VLC (when playing remuxed TS with the Blu-ray sup).
But I'm sure that the original TS is correctly displayed by MPC-HC and VLC because they display the same subtitle size than my DVB recorder.
That's why I'm sure SE reduce the size of these DVB-SUB, but why ... :confused:
3)when I make OCR with Tesseract, a lot of words are not correctly identified while they are well readable (but maybe a little bit too tiny for Tesseract)
4)even when they are correctly identified, I have to skip them all one by one because SE don't know these words while they are very common :o
Nikse555
3rd October 2014, 14:18
@Mole: In next update the OCR fix list and the names etc list will have a "user" file.
@minhjirachi: thx, should be fixed here: http://www.nikse.dk/SubtitleEdit.zip (on src on github + next release)
@Music Fan:
1) the crash should be fixed here: http://www.nikse.dk/SubtitleEdit.zip (on src on github + next release)
Could you verify?
2) Could I check the file (email/msg)? Perhaps VLC/MPC-HC scales the subs etc?
3) yes, small fonts don't work too well in Tesseract.
4) English vs US? Try another dictionary.
Music Fan
4th October 2014, 13:43
Thanks Nikse for all responses.
@Music Fan:
1) the crash should be fixed here: http://www.nikse.dk/SubtitleEdit.zip (on src on github + next release)
Could you verify?
That's a little bit better : files that crashed yesterday can now be parsed but the crash happens after : when I export in Blu-ray sup (without OCR) or when launching OCR.:o
2) Could I check the file (email/msg)?
Here is a 1080i file that can be opened and "OCRed" ;
http://www72.zippyshare.com/v/94029699/file.html
You should be able to play it in VLC and MPC-HC to see the subtitles size (activate the french sub if not done automatically).
When exported in 1080p Blu-ray sup (whatever the bottom line setting) and remuxed in TS with TSmuxer, the subtitles size is much more little.
By the way, could you add the possibility to avoid the bottom line setting to have exactly the same position than in the original sub (if you solve the resize problem, because currently the position changes anyway because of the resize) ?
Perhaps VLC/MPC-HC scales the subs etc?
I don't believe so, they are displayed at the same size than with my DVB recorder.
3) yes, small fonts don't work too well in Tesseract.
Mmmh, could the other method work ?
4) English vs US? Try another dictionary.
That's french, I didn't realize I had to change the language :rolleyes: :D, that's ok now (I tried with the old 3.4.2 version, not the one you linked yesterday) :)
But a lot of t and q are considered as (
:confused:
And I had to download the french dictionnary while there was already a fr-FR.xml file in the language folder. Or does it only concern the language of the program ? If yes, where are stored the dictionnaries ?
Anyway, there are still two problems ; crash with some ts and resize in Blu-ray sup.
Music Fan
5th October 2014, 20:21
That's french, I didn't realize I had to change the language :rolleyes: :D, that's ok now (I tried with the old 3.4.2 version, not the one you linked yesterday) :)
But a lot of t and q are considered as (
:confused:
I found a trick that seems to solve this problem due to little characters size : after exporting in Blu-ray sup with SE (without OCR), I resize the sup with BDSup2Sub ("scale" setting). The text becomes bigger inside the sup which keeps its resolution (1080p sup remains in 1080p).
I chose 2 for x and y (twice the size for width and height to keep text's proportions).
Then I open this sup in SE and launch the OCR which is better done :)
Nikse555
5th October 2014, 20:42
@Music Fan: OK, nice - in the meantime I've hopefully fixed the last TS crash here: http://www.nikse.dk/SubtitleEdit.zip (also in latest src on GitHub)
Music Fan
5th October 2014, 21:02
Thanks, it doesn't crash anymore :)
I recall my other questions ;
1)did you try my ts file to see why it's resized when converted in Blu-ray sup ?
2)could you add the possibility to avoid the bottom line setting to have exactly the same position than in the original sub ?
3)in which folder are stored the dictionnaries ?
;)
minhjirachi
6th October 2014, 03:29
Can you add the subtitle preview, which like as BDSup2Sub? I think it's better when re-check all the subtitle to correct position. With current preview, I can't see nothing. Just the color, style without pos.
mantis2892
12th October 2014, 20:00
I have a problem when converting from SUP to SRT.
When there are two sets of subtitles shown, one at the top of the screen and one at the bottom at the same time, the timings of the bottom get screwed up :(
I mean lets say a subtitle appears at the top of the screen from 15:30:20 - 18:20:12, and lets say something appears at the bottom from 17:15:15 - 19:10:12. With the OCR method, the SRT file generates timecodes of 15:30:20 - 18:20:12 for the top subtitle and 18:20:12 - 19:10:12 for the bottom one.
Is there anyway to fix this (except for manually setting the timecodes of course).
minhjirachi
13th October 2014, 02:25
I have a alignment problem with this script:
Phim thuyết minh độc quyền tại
<font color=yellow><b> WWW.ABC.NET </b></font>
And with the .ass subtitle, when I export to .sup, it has lost everything like subtitle position, color, and the subtitle layers.
Still have error with time stamp when exporting as xml/png.
Subtitle Edit cannot re-load the previous setting.
Music Fan
21st October 2014, 14:01
v3.4.3 has been released 8 days ago, I don't understand why Nikse555 didn't tell us this ;
https://github.com/SubtitleEdit/subtitleedit/releases
Nikse555
21st October 2014, 19:32
He, I only updated the title of this thread.
SE 3.4.3 has some general performance improvements + some fixed memory leaks (especially in batch convert).
Also, "Compare" + "Measurement converter" are broken... not on purpose though ;)
Music Fan
21st October 2014, 22:00
1)did you try my ts file to see why it's resized when converted in Blu-ray sup ?
2)could you add the possibility to avoid the bottom line setting to have exactly the same position than in the original sub ?
3)in which folder are stored the dictionnaries ?
I answer my 3rd question ; on Windows 7, the dictionnaries for Tesseract are stored in ;
C:\Users\user name\AppData\Roaming\Subtitle Edit\Tesseract\tessdata
I needed to know this because as it is optional and installed by download when SE is running, I had to copy this folder for a pc not connected to the web.
Nikse, could you answer the 2 first questions please ? You asked me a TS file a few weeks ago, I guess you tried it.
Thunderbolt8
2nd November 2014, 13:55
again regarding .ass subs and hearing impaired removal, could you please implement some more stuff:
I'm talking about such lines as e.g.
1. {\an4\pos(691,748)}(Chuckles) Yes, ok.
2. {\an4\pos(691,748)}(radio noise)
3. {\an4\pos(691,748)}SPEAKER: text blabla.
1. if a line begins with {\ could you then please implement that "remove text between" '(' and ')' only affects that part of the line which comes after the closing bracket } ? otherwise subtitle screen position information gets deleted as well. (atm the result looks like this: {\an4\pos}Yes, ok. )
1 & 2. and also, if a line begins with {\ and is completely empty after the closing bracket } could you please change it that only such a line gets removed entirely when remove text between '{' and '}' is ticked and no such lines which still include information after the closing bracket } ?
3. and also that "remove text before a colon" does not remove text before a colon which is included in between the {\ and } part of a line. (atm the result looks like this: text blabla)
thanks!
edit: in short, ultimately its about excluding the {\an4\pos(691,748)} stuff in .ass subs from HSD removal.
Nikse555
2nd November 2014, 21:31
@minhjirachi:
I've tried to fix the center issue here: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
What settings is not re-loaded in the above version?
Only some settings works for ass/ssa... font size, color, name etc should be used from the style.
@Music fan:
1) I cannot find other subtitles... tried a few different tools. I do believe that it's resized.
2) Not sure where you mean?
3) You can also use the menu: Spell check -> Get dictionaries -> Open dictionary folder
@Thunderbolt8: could you give some examples of the most important issue regarding remove text for HI?
Thunderbolt8
2nd November 2014, 23:42
@Nikse555 check my last posts up to august 18th. some of the former stuff might be redundant taking my last post into regard. implementing what I wrote in my last post might be the best idea atm for .ass subs.
Nikse555
3rd November 2014, 18:56
@Thunderbolt8: thx for the clarification - could you test this a bit (just updated - also in latest src on github): http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
Music Fan
3rd November 2014, 22:38
@Music fan:
1) I cannot find other subtitles... tried a few different tools. I do believe that it's resized.
2) Not sure where you mean?
3) You can also use the menu: Spell check -> Get dictionaries -> Open dictionary folder
1)Resized by players ? I don't believe so, I explained why.
Anywyay, do you know why DVB-SUB are resized by SE when converted in Blu-ray sup without OCR ?
Other possibility : DVB-SUB may have a real size different than its display size which would be stored in the header of the sup track (or may be linked to the video resolution) ; if SE ignores this header, that could explain why the sup appears more little when converted in Blu-ray sup (which would be its real size).
Otherwise, I don't understand why it becomes more little only when converted in Blu-ray sup with SE.
2)in the Blu-ray sup window, when avoiding OCR (right click in OCR window), otherwise there is a little position change and one has to test several bottom line settings to find the original subtitles position.
3)I was talking about the dictionaries used for OCR, but I found where they are located (C:\Users\user name\AppData\Roaming\Subtitle Edit\Tesseract\tessdata).
Thunderbolt8
3rd November 2014, 23:13
from what I can see that last implementation looks good so far. theres no wrong deletion of positional information within the {\ } brackets any more :)
two more things I noticed though:
1. deletion of residual hyphens '-' could still be improved (refers to both .ass and .srt subs). sometimes after SHD removal, a single or two hyphens remain (as only parts of a line), this should then be deleted as well. e.g.
{\an4\pos(782,897)}<i>- [Grunting Continues]</i> --> {\an4\pos(782,897)}<i>-</i>
works correctly:
- Hello? --> - Hello?
- [ Man ] Paul? --> - Paul?
doesnt work correctly:
- I insist.
<i>- [ Woman Laughing]</i> --> - I insist.
- Oh.
<i>- Yeah.</i> --> - <i>- Yeah.</i>
guess this might get tricky because you'd need to prevent the deletion of hyphens which are part of speech or a sentence in contrast to the deletion of hyphens as indicator of change of the speaker. could be hard in case of single lines. maybe in such a case it would be possible to distinguish between hyphens which are there from the start and such which occured only after the SHD removal process? dunno. but maybe all this could still work out without breaking the other thing :p
2. partly also refers back to the first one, is that in .ass subs lines which have the same timestamps should be regarded as belonging together in terms of SHD removal. e.g.
00:07:04,130 00:07:08,760 {\an4\pos(643,838)}[BOTH SPEAKING IN
00:07:04,130 00:07:08,760 {\an4\pos(643,904)}FOREIGN LANGUAGE]
or
00:07:08,930 00:07:12,760 {\an4\pos(521,772)}(AFU-RA FEATURING JAHDON
00:07:08,930 00:07:12,760 {\an4\pos(519,832)}& KARDINAL OFFISHALL'S
00:07:08,930 00:07:12,760 {\an4\pos(519,904)}"DEAL WIT IT" PLAYING)
get ignored because the brackets () [] dont close at the same line, but those lines are nevertheless displayed as belonging together on the screen at the same time
00:02:56,170 00:02:58,800 {\an4\pos(970,772)}MAN 1: Morning, staff sergeant.
00:02:56,170 00:02:58,800 {\an4\pos(970,653)}Morning, General.
00:02:56,170 00:02:58,800 {\an4\pos(106,905)}MAN 2: Morning, sergeant.
gets changed to:
00:02:56,170 00:02:58,800 {\an4\pos(970,772)}Morning, staff sergeant.
00:02:56,170 00:02:58,800 {\an4\pos(970,653)}Morning, General.
00:02:56,170 00:02:58,800 {\an4\pos(106,905)}Morning, sergeant.
but it should be:
00:02:56,170 00:02:58,800 {\an4\pos(970,772)}- Morning, staff sergeant.
00:02:56,170 00:02:58,800 {\an4\pos(970,653)}Morning, General.
00:02:56,170 00:02:58,800 {\an4\pos(106,905)}- Morning, sergeant.
two different persons are speaking here, but the viewer wouldnt be able to make out that difference on the screen, because adding of the hyphens (as it happens if this occurs within one single line) is missing.
00:41:51,170 00:41:53,050 {\an4\pos(708,827)}- Focus! Focus, Jean!
00:41:51,170 00:41:53,050 {\an4\pos(708,897)}- [ Rattling Stops ]
gets changed to:
00:41:51,170 00:41:53,050 {\an4\pos(708,827)}- Focus! Focus, Jean!
but it should be:
00:41:51,170 00:41:53,050 {\an4\pos(708,827)}Focus! Focus, Jean!
here we have again residual '-' because I guess both lines are not regarded as belonging together and as it seems deletion of hyphens in single lines is tricky because hyphens belonging to actual speech could be affected... notice though how in this example the last line gets deleted entirely and in the very first example at the top the hyphens still remains. could be that the line is not regarded as empty in the top example because of the italic tags (maybe a little bug).
correct SHD removal of subs, especially in case of such .ass subs can get very tricky, I know ;)
Nikse555
4th November 2014, 21:19
@Music Fan:
1) Yeah, there could surely be something SE don't know about TS DVB subtitle pictures! I cannot see anymore from the docs and the java ts parsers. Perhaps some else knows about this!?
(perhaps subtitles are added to the video - before the video player resizes the video to full screen)
The subtitles from TS files are not resized when converted to bluray sup files.
2) Hm, I guess the bottom margin could be calculated (not easy though) - but I guess it could be different for each subtitle?
@Thunderbolt8: I've tried to fix first issue (deletion of residual hyphens - thx for finding it) here: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
Yeah, the multiple lines below each other with same time codes seems advanced...
Music Fan
5th November 2014, 10:05
@Music Fan:
1) Yeah, there could surely be something SE don't know about TS DVB subtitle pictures! I cannot see anymore from the docs and the java ts parsers. Perhaps some else knows about this!?
(perhaps subtitles are added to the video - before the video player resizes the video to full screen)
The subtitles from TS files are not resized when converted to bluray sup files.
2) Hm, I guess the bottom margin could be calculated (not easy though) - but I guess it could be different for each subtitle?
1) As the subtitles size is the same with VLC (and also MPC-HC) than with my DVB recorder, I guess VLC does not resize subtitles (or make the same resize than my DVB recorder). You could maybe look at VLC sources to see how DVB-SUB are handled ?
2) I believed all sup formats had the same resolution than videos (480, 576, 1080 ...) including a lot of transparent lines to be displayed over the video. If it's the case, there should be no need to calculate the bottom margin of original DVB-SUB because this margin is a part of subtitles (I mean lines among others ; some with text, some with only transparency).
When I convert Hd-dvd sup to Blu-ray sup with BDSup2Sub, I don't have to specify bottom margin, I believed it could be as simple for DVB-SUB (when OCR is not needed).
Thunderbolt8
5th November 2014, 17:32
@Thunderbolt8: I've tried to fix first issue (deletion of residual hyphens - thx for finding it) here: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
Yeah, the multiple lines below each other with same time codes seems advanced...will try out the new one.
maybe it help, the multiple line stuff should only apply to .ass subs. afaik not to other formats (well, I only know .srt and .ass :D anyway)
Thunderbolt8
6th November 2014, 15:25
residual hyphen deletion looks good in most cases, but apparently there are still problems with italics:
{\an4\pos(708,897)}- [ Rattling Stops ] <-- was already fine
{\an4\pos(782,897)}<i>- [Grunting Continues]</i> --> {\an4\pos(782,897)}<i>-</i> <-- still not fine
apart from that, I found something else related to it: residual comma deletion after SHD removal :P
Oh, yes. It was, um — --> Yes. It was, —
I ended up — --> I ended up —
its working fine though in this case:
<i>- Mikey!</i> --> <i>- Mikey!</i>
<i>- Aw, man!</i> --> <i>- Man!</i>
Thunderbolt8
8th November 2014, 16:01
residual hyphen deletion looks good in most cases, but apparently there are still problems with italicsanother nice example:
- [panting] <i>Felt that one day,</i> --> - <i>Felt that one day,</i>
Thunderbolt8
14th November 2014, 18:05
and another case of a residual hyphen which should be deleted:
WOMAN: A glass of champagne, please.|- (Laughter) --> - A glass of champagne, please.
works as intended though if there was additionally actual speech information in the 2nd line instead of only SHD stuff.
Thunderbolt8
1st December 2014, 00:31
@Music Fan:
@Thunderbolt8: I've tried to fix first issue (deletion of residual hyphens - thx for finding it) here: http://www.nikse.dk/SubtitleEdit.zip (beta, portable version)
Yeah, the multiple lines below each other with same time codes seems advanced...have there been with the release of 3.4.4 (or with the commits made after) additional changes made compared to this beta above from November 4th? if so then I'd check on those subs again.
Nikse555
1st December 2014, 19:48
@Thunderbolt8: Yes, I've fixed some of the issues in 3.4.4... I hope. Please do test :)
Blueray sup reading will crash on some files (with multi image subs) in 3.4.4 but that is fixed in current c# source on github - and here also: http://www.nikse.dk/SubtitleEdit.zip (portable version, beta)
Some image export issues have also been addressed in latest beta: https://github.com/SubtitleEdit/subtitleedit/issues/343
aax
13th December 2014, 02:02
Hey nikse,
any chance you can add groupings to Multiple Replace tool? With names for each group, and checkmarks for both the whole group and individual actions inside the group.
Music Fan
13th December 2014, 12:23
There is a strange thing with export in sub/idx : the size is not the same than if I export first in Blu-ray sup then convert the sup in sub/idx, while I export both in 720p.
Is there a way to add automatically a space before ? and ! (and maybe also : and ; ) when there is no space between the last word of a sentence and these signs ?
aax
13th December 2014, 20:58
Is there a way to add automatically a space before ? and ! (and maybe also : and ; ) when there is no space between the last word of a sentence and these signs ?
Use replace, set to regular expression:
find(?<! )([:;!?])
replace with
$1
(space before dollar symbol)
You can also add this to "Tools → Multiple replace..." so you can easily run it later.
Music Fan
13th December 2014, 22:20
Thanks but that does not work.
edit : works now, I believe I forgot the space before $ ;) (multiple replace is in edit and not tools, below replace)
By the way, could you explain why you type (?<! )([:;!?]) and not only ([:;!?]) or (:;!?) (without [] ) ?
Betsy25
13th December 2014, 22:50
Thanks but that does not work.
edit : works now, I believe I forgot the space before $ ;) (multiple replace is in edit and not tools, below replace)
By the way, could you explain why you type (?<! )([:;!?]) and not only ([:;!?]) or (:;!?) (without [] ) ?
I think this is a negative lookbehind in regular expressions. the (?<! ) in front literally means "Do not match when there's (already) a space here"
http://www.regular-expressions.info/lookaround.html
The brackets [] mean "any of the following". If they would not be there, it would be seen as the string ":;!?" .
Music Fan
13th December 2014, 22:57
Ok thanks.
Music Fan
14th December 2014, 00:44
Actually I guess I can also write (?<!<:<; )([:;!?]) instead of (?<! )([:;!?]) to avoid to add space before : and ; (and not only ? and !) when there is already a space before these signs, right ?
And do you know how to do it for ... (3 points) ?
Is it something like this for find ;
(... )([...])
and this for replace with ;
$1
?
aax
14th December 2014, 02:48
Check out that Betsy's link and this cheat sheet (http://www.rexegg.com/regex-quickstart.html) for info on regular expressions.
As Betsy already said, the (?<!) is a command that will check if the characters that you put inside it come before whatever else you put next in your search string. Then it will match your search string only if they don't.
(?<!<:<; )([:;!?]) doesn't make sense in your case because it will look for :;!? that aren't precede by "<:<; ".
The expression I gave you in my first reply won't add a space before any of your characters (:;!?) if there is already a space there.
Music Fan
14th December 2014, 03:10
Ok thanks, but when you say "the characters that you put inside it", I guess you mean the characters after ?<!, thus only the space in this case.
For the 3 points, is there a solution with multiple replace or it can't work because it's 3 tree times one character and not a triple character (which doesn't even exist I think) ?
Otherwise I simply go in replace, type ..., search it and add a space before it.
aax
14th December 2014, 03:25
Ok thanks, but when you say "the characters that you put inside it", I guess you mean the characters after ?<!, thus only the space in this case.
Exactly. But check the cheat sheet, there are good examples.
For adding space before three points (a.k.a. ellipsis) when there is none this should do it:
find: (?<![\. ])\.{3}(?!\.)
replace with: space and three dots
In regular expressions dot matches any character except line breaks, so when you want it to mean an actual dot you have to "escape" it by putting a backslash before it.
Music Fan
14th December 2014, 03:51
Thanks for your explanations, I wouldn't have understood it alone, the cheat sheet is not very clear to me ;)
I guess that in your first example, replace with $1 means "let the character, no matter which of the 4, but add a space in front of it", because there are four characters and none is in the replace with line.
aax
14th December 2014, 04:25
Parentheses mark groups that can be used later. You can put them wherever it's convenient. Then you can recall the first group with $1, the second with $2 and so on. (But lookbehind and lookahead don't count as groups because parentheses are a part of their commands, that's why you're replacing with $1 and not $2.)
Music Fan
14th December 2014, 17:53
Ok, interesting.
Thunderbolt8
25th December 2014, 17:43
I Have two files .srt files attached which still show problems with SHD removal. that should make it easier for you to see for yourself first if the changes you made are doing what they are supposed to do http://www.sendspace.com/filegroup/aL1hsk990Hrz6JX%2BwXUNmg
in case of the 'halloween' one there are quite a few lines which should get removed but are not:
? My Paul ?
? I give you all ?
? No keys ?
those files remain completely untouched despite of the remove text between '?' and '?' option. maybe its because of the spaces? anyway, those lines should be removed.
the 2nd one 'safe' is much more tricky. there is still quite some stuff which is not affected correctly by SHD removal (maybe you forgot to add stuff you added for ( ) brackets to affect other types of brackets as well [ ] ?)
<i>- ♪♪[ Upbeat Dance ]</i>
- Big push. Four more.
--> if "remove text if it contains ♪" is ticked both lines get removed even though only the first one and the hyphen of the 2nd line should be. if unticked, the result is:
<i>- </i>
- Big push. Four more.
I'm sorry. I, um — --> I'm sorry. I, —
--> should be: I'm sorry. I— (afaik this already works(?) if there was a regular hyphen - or double hyphen -- instead of that special character for a hyphen)
<i>- a man who wants to make his mark...
- [ Coughing]</i>
--> <i>- a man who wants to make his mark...</i> --> the hyphen needs to be removed.
<i>- Mm-hmm.</i>
- in my spare time.
--> <i>- </i>
- in my spare time.
--> first line and hyphen of 2nd line need to be removed.
there was also a false positive in another subtitle track which was:
- And you?
- I —
--> And you?
should be: untouched, the 2nd line is fine as it is.
Music Fan
25th December 2014, 19:41
I have problems with multiple replace, I don't find the codes to replace some signs or set of signs.
I'dl ike to replace
7
(space and seven)
by
?
because some ? were interpreted as 7
And also
'?
by
?
The apostrophe shouldn't be there.
and
' ?
by
?
Same with space between ' and ?
Thanks.
edit : actually it works if I choose normal instead of regular expression. Some corrections need to choose normal to work, others need regular expression, I still have to understand differences between these settings.
Nikse555
30th December 2014, 14:37
@Thunderbolt8: thx for the info and testing :)
New beta version is here: http://www.nikse.dk/SubtitleEdit.zip
Let me know how it works.
@Music Fan: Regular expressions are not normal strings but more like a programming language... check this tutorial: http://www.codeproject.com/Articles/9099/The-Minute-Regex-Tutorial
Also, you can right-click in the "Find" textbox in SE for some common expressions (they are kinda hard to remember...)
Thunderbolt8
30th December 2014, 23:38
thanks. However, I have a problem accessing the hearing impaired removal screen with the 'safe' subtitle file now after replacing the old files with the new beta ones. but this seems only to affect this specific file as it works fine in case of other ones. I get this unhandled exception message:
See the end of this message for details on invoking
just-in-time (JIT) debugging instead of this dialog box.
************** Exception Text **************
System.ArgumentOutOfRangeException: Index and count must refer to a location within the string.
Parameter name: count
at System.String.Remove(Int32 startIndex, Int32 count)
at Nikse.SubtitleEdit.Logic.Forms.RemoveTextForHI.RemoveTextFromHearImpaired(String text)
at Nikse.SubtitleEdit.Forms.FormRemoveTextForHearImpaired.GeneratePreview()
at Nikse.SubtitleEdit.Forms.Main.RemoveTextForHearImparedToolStripMenuItemClick(Object sender, EventArgs e)
at System.Windows.Forms.ToolStripMenuItem.OnClick(EventArgs e)
at System.Windows.Forms.ToolStripItem.HandleClick(EventArgs e)
at System.Windows.Forms.ToolStripMenuItem.ProcessCmdKey(Message& m, Keys keyData)
at System.Windows.Forms.ToolStripManager.ProcessShortcut(Message& m, Keys shortcut)
at System.Windows.Forms.Form.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.ContainerControl.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.ContainerControl.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.ContainerControl.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.ProcessCmdKey(Message& msg, Keys keyData)
at System.Windows.Forms.Control.PreProcessMessage(Message& msg)
at System.Windows.Forms.Control.PreProcessControlMessageInternal(Control target, Message& msg)
at System.Windows.Forms.Application.ThreadContext.PreTranslateMessage(MSG& msg)
************** Loaded Assemblies **************
mscorlib
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34014 built by: FX45W81RTMGDR
CodeBase: file:///C:/Windows/Microsoft.NET/Framework64/v4.0.30319/mscorlib.dll
----------------------------------------
SubtitleEdit
Assembly Version: 3.4.4.115
Win32 Version: 3.4.4.115
CodeBase: file:///C:/Program%20Files%20(x86)/Subtitle%20Edit/SubtitleEdit.exe
----------------------------------------
System.Windows.Forms
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.33440 built by: FX45W81RTMREL
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Windows.Forms/v4.0_4.0.0.0__b77a5c561934e089/System.Windows.Forms.dll
----------------------------------------
System.Drawing
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.33440 built by: FX45W81RTMREL
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Drawing/v4.0_4.0.0.0__b03f5f7f11d50a3a/System.Drawing.dll
----------------------------------------
System
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34239 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System/v4.0_4.0.0.0__b77a5c561934e089/System.dll
----------------------------------------
System.Xml
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34230 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Xml/v4.0_4.0.0.0__b77a5c561934e089/System.Xml.dll
----------------------------------------
System.Core
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.33440 built by: FX45W81RTMREL
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Core/v4.0_4.0.0.0__b77a5c561934e089/System.Core.dll
----------------------------------------
System.Configuration
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.33440 built by: FX45W81RTMREL
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Configuration/v4.0_4.0.0.0__b03f5f7f11d50a3a/System.Configuration.dll
----------------------------------------
Accessibility
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.33440 built by: FX45W81RTMREL
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/Accessibility/v4.0_4.0.0.0__b03f5f7f11d50a3a/Accessibility.dll
----------------------------------------
************** JIT Debugging **************
To enable just-in-time (JIT) debugging, the .config file for this
application or computer (machine.config) must have the
jitDebugging value set in the system.windows.forms section.
The application must also be compiled with debugging
enabled.
For example:
<configuration>
<system.windows.forms jitDebugging="true" />
</configuration>
When JIT debugging is enabled, any unhandled exception
will be sent to the JIT debugger registered on the computer
rather than be handled by this dialog box.
Nikse555
31st December 2014, 00:45
@Thunderbolt8: thx again - the crash should now be fixed: http://www.nikse.dk/SubtitleEdit.zip
Thunderbolt8
31st December 2014, 15:09
thanks for fixing. I tested the 'safe' subtitle file again and some problems still persist:
line 30:
<i>- ♪♪[ Upbeat Dance ]</i>
- Big push. Four more.
when unticking the "remove text if it contains ♪" box then the result is like this:
<i></i>- Big push. Four more.
the hyphen between the brackets is gone, but the brackets are still there (which itself shouldnt be that bad as it shouldnt show during watching) and the 2nd hyphen is still there as well which needs to be removed (e.g. italic boxes with no information between them should be disregarded as textual information of a line so you would have a single hyphen in a single line which then can get removed)
then the problem with , — still exists:
line 62 (or 363 or others)
- I just, uh —
- What?
--> - I just, —
- What?
the comma and the space after that still need to go. (or at least the comma, the space is debatable. dont know if its possible to implement this but the best solution would be to check whether there always is a space in front of the - or — symbol in other examples in the corresponding subtitle track or not and then correct it following the pattern, either with removing the space or leaving it in. but this might be pretty advanced stuff :) )
line 340
- Oh — Oh, my God!
- Oh, my God.
--> - — My God!
- My God.
this one is really tricky, I dont really know if its possible to fix this without breaking anything else. in this case the 2nd hyphen would have to go because both words in front of and after it will be removed as well (while usually such a hyphen as sign of stuttering etc. shouldnt be removed). guess its not a biggie if this cannot be fixed without breaking anything else.
happy new year :)
Music Fan
31st December 2014, 16:27
the hyphen between the brackets is gone, but the brackets are still there (which itself shouldnt be that bad as it shouldnt show during watching) and the 2nd hyphen is still there as well which needs to be removed (e.g. italic boxes with no information between them should be disregarded as textual information of a line so you would have a single hyphen in a single line which then can get removed)
Did you try to remove it by the replace function ?
Thus you make it in 2 steps ;
1) do what you did (to remove text between the brackets)
2) then replace <i></i>- by - (to remove the brackets)
But I'm not sure it will work, I don't know if brackets can be considered as text.
Thunderbolt8
31st December 2014, 20:05
I already have some stuff added there, but a proper SHD removal would be preferred.
Nikse555
1st January 2015, 01:48
Happy new year :)
@Thunderbolt8: thx - new version up again - http://www.nikse.dk/SubtitleEdit.zip (portable version, beta)
Thunderbolt8
1st January 2015, 13:38
thank you!
quick feedback: example 2 and 3 of post number 294 are fixed, example 1 is still the same as before
Thunderbolt8
2nd January 2015, 16:56
would it please be possible for you to try to work at these SHD fixes again? topic was taking multiple (2-3) lines in .ass (only!) subs into account for SHD removal, because they appear at the screen at exactly the same time and can potentially appear clumbed together so you wouldn't know theres a change of speaker after current SHD removal.
01:08:40.570 1:08:42.660 {\an4\pos(556,723)}MAN1: What's wrong with you?
01:08:40.570 1:08:42.660 {\an4\pos(841,823)}MAN2: You crazy?
01:08:40.570 1:08:42.660 {\an4\pos(838,897)}You outta your mind?
which gets changed to
01:08:40.570 1:08:42.660 {\an4\pos(556,723)}What's wrong with you?
01:08:40.570 1:08:42.660 {\an4\pos(841,823)}You crazy?
01:08:40.570 1:08:42.660 {\an4\pos(838,897)}You outta your mind?
but should be:
01:08:40.570 1:08:42.660 {\an4\pos(556,723)}- What's wrong with you?
01:08:40.570 1:08:42.660 {\an4\pos(841,823)}- You crazy?
01:08:40.570 1:08:42.660 {\an4\pos(838,897)}You outta your mind?
in case this proves to be difficult I guess you could have a look at Subextractor http://subextractor.codeplex.com/ which is able to do this when you save OCRed file as .ass and you tick SHD removal. maybe there is something which can help you implementing this in SubtitleEdit ;)
Bozotheclown
14th January 2015, 23:39
Hi,
Very nice subeditor. Especially for bitmap subs. (read and ocr)
I'am looking for solution to ocr'ed subs from dvb recorded stream. CCextractor - failed, SubEdit - do the job.
I want to ask if preview for original bitmap subs (like in ocr window) is possible during spell checking (an option to load original sub in spellcheck window).
kalehrl
26th January 2015, 13:59
Hi Nikse
Is there a way to join 2 subtitles so that they will be completely merged.
There is an option 'join subtitles' but it actually just appends one subtitle to the end of other.
They are not completely merged.
StainlessS
28th January 2015, 00:11
They are not completely merged.
You might want to define "merged".
kalehrl
28th January 2015, 13:04
You might want to define "merged".
I expected the second subtitle to be included respecting its time codes, not to be appended at the end of another subtitle.
Please have a look at the picture:
http://i60.tinypic.com/15zrprb.png
The first entry in the second subtitle has start time of 00:14:30,002. Yet, it is shown after another line with end time of 00:57:48,299.
Music Fan
28th January 2015, 15:17
You can try to change timecodes of 2nd subtitle file before to merge both files (go to synchronization, adjust all times).
If it comes from a split movie, I guess you should add the duration of the first video part, not the timecode of the first subtitle file's last line, because the first video probably ends a few seconds or minutes after its last text's timecode.
foxyshadis
29th January 2015, 00:06
You can try to change timecodes of 2nd subtitle file before to merge both files (go to synchronization, adjust all times).
If it comes from a split movie, I guess you should add the duration of the first video part, not the timecode of the first subtitle file's last line, because the first video probably ends a few seconds or minutes after its last text's timecode.
That's easy, you only need to do two things:
Tools->Sort by->Start time
Tools->Renumber
Music Fan
29th January 2015, 21:48
Tools->Sort by->Start time
Tools->Renumber
This doesn't make any effect on timecodes.
And anyway, that's not enough to merge files.
StainlessS
4th February 2015, 01:13
Thanks for the update Nikse :thanks:
Thunderbolt8
12th February 2015, 20:48
would it be possible to get a new beta version?
Thunderbolt8
14th February 2015, 21:41
I tried to compile my own version from the master.zip, but ran into problems. I installed visual studio 2013 and the git command line tool, but when I run build.bat then I get
"Could not run Git - build number will be 9999!"
and
"Inno Setup wasn't found; the installer wasn't build."
whats going on here? (please take into consideration I have no clue about programming, compiling etc.)
Thunderbolt8
17th February 2015, 12:34
I tried to compile my own version from the master.zip, but ran into problems. I installed visual studio 2013 and the git command line tool, but when I run build.bat then I get
"Could not run Git - build number will be 9999!"
and
"Inno Setup wasn't found; the installer wasn't build."
whats going on here? (please take into consideration I have no clue about programming, compiling etc.)welp, anyone?
jpsdr
18th February 2015, 09:56
For git, i don't know, but maybe for the "Inno Setup wasn't found; the installer wasn't build." message is because you don't have the programme for building the installer installed (as far as i know, VS2013 don't build/create installer).
http://www.exemsi.com/inno-setup-and-msi
Betsy25
18th February 2015, 11:40
welp, anyone?
I think you should have Git, 7-zip, Inno Setup 5 and VS2013 installed.
Git (for windows) : https://windows.github.com/ (don't mind the steps explained afterwards, just simply install so there's an environment path set on your system)
7-zip : http://www.7-zip.org/
Inno Setup 5 : http://www.jrsoftware.org/isdl.php (download & install the UNICODE version)
Music Fan
18th February 2015, 13:05
Is there a way to remove empty lines (I mean timecodes without text) and renumber srt's lines ? Of course I can remove empty lines manually in the srt then use the renumber function in SE but a "remove empty lines" function would be easier.
Currently, the renumber function does not remove timecodes without text.
Nikse555
18th February 2015, 14:13
@Thunderbolt8: What happens if you open "SubtitleEdit.sln" in Visual Studio and press Ctrl+Shift+b (build) ?
"build.bat" builds the complete installer which you might not need.
@Music Fan: To remove empty lines, you can use Tools -> Fix common errors - select first fix action.
SE also has an option to remove empty lines when opening subtitles.
Music Fan
18th February 2015, 15:54
Thanks ;)
Thunderbolt8
18th February 2015, 16:29
@Thunderbolt8: What happens if you open "SubtitleEdit.sln" in Visual Studio and press Ctrl+Shift+b (build) ?that seems to work, I get a Build: 4 succeeded message.
But I dont really know where the compiled file have been output. are they placed in the \src\bin\debug directory?
Nikse555
18th February 2015, 20:24
But I dont really know where the compiled file have been output. are they placed in the \src\bin\debug directory?
Yes :)
And if your change the drop-down-box in the toolbar from "Debug" to "Release" the exe file will be in \src\bin\Release.
Thunderbolt8
18th February 2015, 21:44
thanks for fixing. I tested the 'safe' subtitle file again and some problems still persist:
line 30:
<i>- ♪♪[ Upbeat Dance ]</i>
- Big push. Four more.
when unticking the "remove text if it contains ♪" box then the result is like this:
<i></i>- Big push. Four more.
the hyphen between the brackets is gone, but the brackets are still there (which itself shouldnt be that bad as it shouldnt show during watching) and the 2nd hyphen is still there as well which needs to be removed (e.g. italic boxes with no information between them should be disregarded as textual information of a line so you would have a single hyphen in a single line which then can get removed)
then the problem with , — still exists:
line 62 (or 363 or others)
- I just, uh —
- What?
--> - I just, —
- What?
the comma and the space after that still need to go. (or at least the comma, the space is debatable. dont know if its possible to implement this but the best solution would be to check whether there always is a space in front of the - or — symbol in other examples in the corresponding subtitle track or not and then correct it following the pattern, either with removing the space or leaving it in. but this might be pretty advanced stuff :) )
line 340
- Oh — Oh, my God!
- Oh, my God.
--> - — My God!
- My God.
this one is really tricky, I dont really know if its possible to fix this without breaking anything else. in this case the 2nd hyphen would have to go because both words in front of and after it will be removed as well (while usually such a hyphen as sign of stuttering etc. shouldnt be removed). guess its not a biggie if this cannot be fixed without breaking anything elseproblem #2 & #3 have been fixed with the latest problems. :)
Problem #1 does still exist though.
theres also one problem I might have overlooked before: Line 999
999
01:14:49,570 --> 01:14:52,573
- My deepest welcome|to Carol and Wade.|<i>- [ Claire ] Ward.</i>
gets changed to
- My deepest welcome to Carol and Wade.
- <i>- Ward.</i>
this one doesnt seem to be a problem if the set of lines consists only of 2 lines. but apparently it doesnt work in case of 3 lines yet.
Keep up the good work! ;)
Thunderbolt8
22nd February 2015, 12:32
found some more cases which need residual hyphens be removed in this file: https://www.sendspace.com/file/senjh6
- What?|- Uh —
gets changed to
What?- —
John. Mr. Malkovich, sir. Uh, um —
gets changed to
John. Mr. Malkovich, sir. —
Well, boy, I'm — Uh —
gets changed to
Well, boy, I'm — —
Thunderbolt8
27th February 2015, 22:43
just build the latest beta and seems like there is a bug. opened a .srt file and got a crash in the hearing impaired removal menu (crtl+shift+h) when I ticked/unticked the remove text before colon but only if text is uppercase box and then everytime I try to enter the hearing impaired removal menu again.
only happens with a certain .srt file though, others are fine.
See the end of this message for details on invoking
just-in-time (JIT) debugging instead of this dialog box.
************** Exception Text **************
System.IndexOutOfRangeException: Index was outside the bounds of the array.
at Nikse.SubtitleEdit.Logic.Forms.RemoveTextForHI.RemoveColon(String text) in c:\Users\Daniel\Desktop\subtitleedit-master\subtitleedit-master\src\Logic\Forms\RemoveTextForHI.cs:line 109
at Nikse.SubtitleEdit.Logic.Forms.RemoveTextForHI.RemoveTextFromHearImpaired(String text) in c:\Users\Daniel\Desktop\subtitleedit-master\subtitleedit-master\src\Logic\Forms\RemoveTextForHI.cs:line 354
at Nikse.SubtitleEdit.Forms.FormRemoveTextForHearImpaired.GeneratePreview() in c:\Users\Daniel\Desktop\subtitleedit-master\subtitleedit-master\src\Forms\RemoveTextFromHearImpaired.cs:line 117
at Nikse.SubtitleEdit.Forms.FormRemoveTextForHearImpaired.CheckBoxRemoveTextBetweenCheckedChanged(Object sender, EventArgs e) in c:\Users\Daniel\Desktop\subtitleedit-master\subtitleedit-master\src\Forms\RemoveTextFromHearImpaired.cs:line 196
at System.Windows.Forms.CheckBox.set_CheckState(CheckState value)
at System.Windows.Forms.CheckBox.OnClick(EventArgs e)
at System.Windows.Forms.CheckBox.OnMouseUp(MouseEventArgs mevent)
at System.Windows.Forms.Control.WmMouseUp(Message& m, MouseButtons button, Int32 clicks)
at System.Windows.Forms.Control.WndProc(Message& m)
at System.Windows.Forms.ButtonBase.WndProc(Message& m)
at System.Windows.Forms.NativeWindow.Callback(IntPtr hWnd, Int32 msg, IntPtr wparam, IntPtr lparam)
************** Loaded Assemblies **************
mscorlib
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34209 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.NET/Framework64/v4.0.30319/mscorlib.dll
----------------------------------------
SubtitleEdit
Assembly Version: 3.4.5.9999
Win32 Version: 3.4.5.9999
CodeBase: file:///C:/Program%20Files%20(x86)/Subtitle%20Edit/SubtitleEdit.exe
----------------------------------------
System.Windows.Forms
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34209 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Windows.Forms/v4.0_4.0.0.0__b77a5c561934e089/System.Windows.Forms.dll
----------------------------------------
System.Drawing
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34209 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Drawing/v4.0_4.0.0.0__b03f5f7f11d50a3a/System.Drawing.dll
----------------------------------------
System
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34239 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System/v4.0_4.0.0.0__b77a5c561934e089/System.dll
----------------------------------------
System.Core
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34209 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Core/v4.0_4.0.0.0__b77a5c561934e089/System.Core.dll
----------------------------------------
System.Configuration
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34209 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Configuration/v4.0_4.0.0.0__b03f5f7f11d50a3a/System.Configuration.dll
----------------------------------------
System.Xml
Assembly Version: 4.0.0.0
Win32 Version: 4.0.30319.34230 built by: FX452RTMGDR
CodeBase: file:///C:/Windows/Microsoft.Net/assembly/GAC_MSIL/System.Xml/v4.0_4.0.0.0__b77a5c561934e089/System.Xml.dll
----------------------------------------
************** JIT Debugging **************
To enable just-in-time (JIT) debugging, the .config file for this
application or computer (machine.config) must have the
jitDebugging value set in the system.windows.forms section.
The application must also be compiled with debugging
enabled.
For example:
<configuration>
<system.windows.forms jitDebugging="true" />
</configuration>
When JIT debugging is enabled, any unhandled exception
will be sent to the JIT debugger registered on the computer
rather than be handled by this dialog box.
edit: fixed, thanks.
Music Fan
2nd March 2015, 11:06
Exactly. But check the cheat sheet, there are good examples.
For adding space before three points (a.k.a. ellipsis) when there is none this should do it:
find: (?<![\. ])\.{3}(?!\.)
replace with: space and three dots
In regular expressions dot matches any character except line breaks, so when you want it to mean an actual dot you have to "escape" it by putting a backslash before it.
Hi,
I have nearly the same question than a few months ago : how to add space after (and not before) 3 points only when there is a letter after these 3 points (and not when the 3 points are the end of the sentence) ?
For example ;
...and he went there.
would become ;
... and he went there.
Betsy25
2nd March 2015, 13:47
Hi,
I have nearly the same question than a few months ago : how to add space after (and not before) 3 points only when there is a letter after these 3 points (and not when the 3 points are the end of the sentence) ?
For example ;
...and he went there.
would become ;
... and he went there.
You can try regex's here : https://regex101.com/
Function should be a Lookahead after the match (http://www.rexegg.com/regex-disambiguation.html#lookahead) (in human terms : "Look if there's something specific after that what you'll want to match")
PS: A good tip which makes it a lot less difficult to decypher, whatever is inside (?somestuffhere) cases will always be checked for, but will never end up being included in the final match string.
Thus, this should do it :
Find : (3 dots, only if they are directly followed by either a letter or a number)
\.{3}(?=[a-z0-9])
Replace with 3 dots + space
Thunderbolt8
2nd March 2015, 19:18
apparently HD DVD subtitles are not recognized and cannot be opened. could you add support for that please?
https://www.sendspace.com/file/5sije6
thanks!
Thunderbolt8
5th March 2015, 21:24
looking at this https://github.com/SubtitleEdit/subtitleedit/commit/eb227e634788d627342a19f813a656373f305324 a question:
does this only fix the problem in this specific example mentioned here? or would this also fix similar cases, e.g. just with different words than used in the example?
another thing, would it be possible to add a kind of ignore list we can edit, consisting of certain phrases or sentences we can exclude from the hearing impaired removal?
for example, I have "oh" in my hearing impaired removal list, but I like to keep it in case of "Oh, my god" or "Oh, dear". I always have to check the complete list of changes manually for that and untick it. I could add the phrase "my god" to the multiple replace list and set it to replace with "oh, my god", but there are sometimes also lines which really are just "my god", which would then changed wrongly to "oh, my god".
perhaps it sounds a little bit picky, but such a kind of ignore removal list could save the work to manually check the entire proposed hearing impaired removal list of a subtitle file before I can apply the changes.
Also, would it be possible to add a match case box we could tick for each single entry of the hearing impaired removal list? for example, I have "er" in that list to remove things like "Er... I dont know" or "wait...er..." or something like that. But then the word "ER" (emergency room) also gets to be removed as a false positive.
Thunderbolt8
6th March 2015, 18:23
there seem to be some bugs with the latest update. sometimes, the SHD removal or fixes in the fix common error menu are not applied. I can press OK, but the relating entries dont get removed and when I open the window again they are all still listed there. doesnt seem to happen always or with all files, though.
e.g. in this one: https://www.sendspace.com/file/8q11ky
Music Fan
6th March 2015, 22:29
Thus, this should do it :
Find : (3 dots, only if they are directly followed by either a letter or a number)
\.{3}(?=[a-z0-9])
Replace with 3 dots + space
Thanks, this works well ;)
apparently HD DVD subtitles are not recognized and cannot be opened. could you add support for that please?
I already asked it, but IIRC Nikse said it was too much work, and as BDSup2Sub supports this conversion, it's less useful to add it in SE ;
http://forum.doom9.org/showthread.php?p=1683707#post1683707
But you can use SE after BDSup2Sub's conversion (in Blu-ray sup) if you need OCR.
Thunderbolt8
6th March 2015, 23:02
there seem to be some bugs with the latest update. sometimes, the SHD removal or fixes in the fix common error menu are not applied. I can press OK, but the relating entries dont get removed and when I open the window again they are all still listed there. doesnt seem to happen always or with all files, though.
e.g. in this one: -https://www.sendspace.com/file/8q11kyit seems the removal of lines consisting only of SHD information like stuff in () or [] or lines consisting just of "aahh" "uh-huh" etc. is broken. those lines dont get removed, no matter how often I click it.
Thunderbolt8
7th March 2015, 11:55
theres still a bug with in the SHD removal process despite the latest fixes. a line with SHD information and regular speech e.g. "[coughs] Well, that it was I mean" now gets entirely removed instead of just the SHD part in the brackets. Its shown correctly in the SHD removal preview window though, just not done correctly during the process.
Nikse555
7th March 2015, 12:19
@Thunderbolt8: thx for the bug reports :)
Latest beta (http://www.nikse.dk/SubtitleEdit.zip - 3.4.5 build 431) seems to work, right?
Edit: He, found the bug... latest beta (http://www.nikse.dk/SubtitleEdit.zip - 3.4.5 build 436) seems to work, right?
Thunderbolt8
7th March 2015, 15:19
yes, seems like it. thanks for the fix! In case Ill find more I report back.
Thunderbolt8
7th March 2015, 22:24
some kind of special case for SHD removal: maybe you could add some exception to the "remove text before a locon (':') only if text is UPPERCASE" scenario in combination with certain names like McWHATEVER. even though the letter c in the name is not uppercase, the name is basically meant to be.
Thunderbolt8
8th March 2015, 23:32
could you please change that the window size of the hearing impaired removal window (ctrl+shift+h) stays the same the way you changed it to last time? This works for the Fix common errors window (ctrl+shift+F), but not the first one.
and some small SHD fix:
Which one do you want to be?|Uh--
gets changed to
Which one do you want to be?--
afaik this also works in case if there is only one hyphen "-" or maybe also the longer special hyphen character, but apparently not yet in the case of double hyphens "--"
Also please add the double hyphen "--" in general to all other sort of stuff which gets checked and corrected when it comes to hyphens and for which the single hyphen and the special hyphen character are already affected, if the double hyphen is not included yet in such hyphen type of scenarios (e.g. stuff related to the fix common errors section. afaiks the "fix first letter to uppercase after paragraph detects a single hyphen (maybe also the special hyphen character) at the end of a line, with no hyphen following in the next line, and does correctly not suggest to capitalize the first letter of that next line (because the sentence is supposed to continue). that doesnt work in case of a double hyphen, though)
some more:
- Mr. Harding?|-Mm-hm. Oh.
gets changed to
- Mr. Harding?
and similarly:
Oh.|-I'm awfully tired.
gets changed to
-I'm awfully tired.
And:
-Sit down. Sit down.|-Oh! Oh!
gets changed to
- Sit down. Sit down.!
Thunderbolt8
9th March 2015, 01:19
another kind of problem/bug: when there is an unneeded period, e.g. "!." then after this instance all following lines or the entire subtitle file, with normal periods are wrongly detected as unneeded periods as well. although theres nothing wrong with those lines.
heres an example: https://www.sendspace.com/file/zv9znz
the undeeded period is in line 79 and after that there occurs the problem (theres also a 2nd one in line 180)
also, in case of .ass subs, there is no missing space in case of }" because the } is part of the line position information on the screen which means that the " is actually the first real character of a line referring to actual speech information. e.g. as in
{\an4\pos(717,888)}"How cool,
so there doesnt need to be a missing space detected in the fix common error section in such a case.
Nikse555
9th March 2015, 21:35
another kind of problem/bug: when there is an unneeded period, e.g. "!." then after this instance all following lines or the entire subtitle file...
Nice catch :)
thx for the bug reports - new portable beta up: http://www.nikse.dk/SubtitleEdit.zip (also on GitHub)
Thunderbolt8
9th March 2015, 23:24
thanks for the quick fixes!
Thunderbolt8
9th March 2015, 23:35
another thing, in some cases when OCRing a subtitle file it can happen that the baseline of dots, commas and exclamation marks is off. meaning they are then put into a seperate line, taken apart from the rest of the speech line.
e.g.
You forgot "cursed"|.
gets correctly changed to
You forgot "cursed".
in the above case with a dot "." this gets detected in the hearing impaired window and the line break is correctly removed. however, that isnt the case for commas or exclamation marks which are off in the same way. its not exactly a part which belongs to SHD removal, but it would be nice if this could get detected & fixed as well, either in the SHD removal window or the fix common error one.
another version of this btw. is when these characters are put into a separate line before the line they belong to:
e.g.
?|How was your day
gets changed to
How was your day ---> should be: How was your day?
(the way this is handled atm is simply to remove the question mark and the line break instead of repositioning the question mark; in case of commas afaik nothing is detected atm; not sure about exclamation marks).
referring to the above, if there is such a misplaced dot then the SHD removal can get a bit screwed up as well:
- and I'll speak to you later|.|- [ Anna] OK.
gets changed to
- and I'll speak to you later .|- - OK. (the space between the dot and the last word gets inserted there for some reason, even though this wouldnt be the case if the 3rd line wasnt there, as seen in the first example at the beginning)
its working correctly though if the dot is put correctly where it belongs, so not sure if "fixing" this is maybe taking it too far.
JayJayH
7th April 2015, 12:49
This is my first post, so please bear with me.
I currently use SubtitleEdit 3.3.15 to edit .SSA subtitles and their timings. On occasions I need two subtitles to be displayed on screen at the same time, for example in red colour at the top of the screen (maybe to display a cell-phones text message) while at the same time another line at the bottom of the screen in green colour containing what is being said. The attached .jpg best illustrates what I mean. Note that after I make changes, I save (ctrl+S) the .SSA file in SE edit and the corrected results are displayed.
I have been unable to achieve the same result in any version of SE later than 3.3.15 no matter what I do - even if I change nothing.
I'm using Win 7 with latest updates.
Any ideas anyone, please?
jpsdr
10th April 2015, 19:59
I've tried to use Subtitle Edit 3.4.5 to OCR on an idx/sub file, with image compare method. Almost each time there is 2 lines of text, it ask me a caracter but considering letter of both lines being one caracter. Is there a parameter to set to avoid this issue ?
Thunderbolt8
12th April 2015, 20:19
some more SHD removal fixes:
WOMAN: <i>Mr. Sportello?</i>|- Mm-hm.
gets changed to:
<i>- Mr. Sportello?</i>
but should get changed to: <i>Mr. Sportello?</i>
--> the WOMAN: speaker information part is not taken into consideration for SHD removal in combination with the hyphen from the speech part after the line break.
minhjirachi
14th April 2015, 14:30
Still using version 3.4.2. The newest version have a little bug when exporting as .sup file or bdn/xml. When exporting as .sup, the program close suddenly. And with the bdn/xml, I can't import to Scenarist.
Nikse555
15th April 2015, 20:50
@JayJayH: Sorry, SE only supports a simple preview of the subtitle (no position + only one sub).
@jpsdr: You could try editing Settings.xml and change "ShowBetaStuff" to true - and then try the "New image compare" ocr...
@Thunderbolt8: thx, the "remove text for HI" issue should be fixed in latest beta: http://www.nikse.dk/SubtitleEdit.zip (portable version)
@minhjirachi: Cannot re-create .sup creating crash - what OS are you using?
I don't have Scenarist... any idea why it does not work?
Also, SE 3.4.6 is out (well, 18 days ago)
jpsdr
16th April 2015, 08:45
Is there a way to get the french dictionnary ? The PC where i'm using SE is totaly offline, and will never be connected.
Betsy25
16th April 2015, 16:02
Is there a way to get the french dictionnary ? The PC where i'm using SE is totaly offline, and will never be connected.
You could try the local library, but I don't know if that one will be of much help here.
Otherwise, while you might be connected, Spell Check / Get Dictionaries... / French / Download.
You can then copy them on a USB and put them in the dictionary folder on the unconnected PC, "Open dictionaries folder..." copy/paste the "fra*.*" and "hyph_fr*.*" files.
Nikse555
16th April 2015, 18:27
Yes, do like Betsy25 said - or do it manually:
1) find .oxt file from libre office (or open office), eg. http://extensions.libreoffice.org/extension-center/dictionnaires-francais/releases/5.3/lo-oo-ressources-linguistiques-fr-v5-3.oxt
2) Download the file and rename the .oxt file to .zip
3) Unzip the file
4) copy the contents of the "dictionaries" folder to SE's "Dictionaries" folder.
jpsdr
17th April 2015, 09:23
Finaly, i've installed SE on a PC with connection, downloaded langage i needed for both things (Teressac & dictionnary) and copied it on USB stick and put them on the other PC. Gladly, it worked. I was afraid there could be some things set in registers or others, but apparently not.
Thanks Nikse555 for this very usefull tools, handling a lot of formats, and i've tested Teressac with the correct langage on sub/idx, results are very good, even excellent.
:thanks:
JayJayH
17th April 2015, 11:06
Hi Nikse555, Thanks for your reply at post #341.
I guess what you're saying is that all versions of SE which are later than v3.3.15 no longer support previews of .SSA subtitles showing where they will finally display.
The screen-shots (taken from SE v3.3.15) attached to my post #337 clearly show that it used to happen with .SSA subtitles in that (and the earlier versions if I remember correctly) version - please see the (upper) RED and (lower) GREEN subtitles.
It seems that feature is now "lost" in the later versions - all that they now display is the WHITE subtitle preview (again, see the screen-shot)
minhjirachi
18th April 2015, 15:19
@Nikse555: I'm using Windows 8.1 64-bit. I have known why the Subtitle Edit got crash. Because it meet the too long line. With the old version (e.g 3.4.2), it forced export to sub and show up the lines, which are too long.
With the Scenarist, please view this picture:
http://i.imgbox.com/fmWIhUAq.png
ndjamena
19th April 2015, 14:24
I've been bashing away at a little program to convert 608 captions taken from m4vs to srt. I'd like to upgrade to Advanced SubStation Alpha, but given the lack of background colours and positioning in the srt format I've pretty much done everything I can for now (other than Flash On and alternate channels). The code is a complete shambles (I was learning as I went) and completely unoptimised but on comparing the SRTs played on VLC to the original captions played through iTunes they're pretty much exactly the same.
Then I discovered Subtitle Edit could load 608 from an m4v. Sadly, it seems to be rather inferior to my programs abilities at the moment, hopefully that can be fixed though.
This is the first few lines as Subtitle Edit extracts it:
1
00:00:07,007 --> 00:00:09,041
HI, THERE. JUST WANTED
TO WELCOME YOU TO MY SHOWSTARRING ME--KUZCO.
2
00:00:09,042 --> 00:00:11,844
SO, NO CHANGING
THE CHANNEL. UNDERSTAND?
3
00:00:11,845 --> 00:00:13,212
NO CHANGING.
OK. THEME MUSIC.
4
00:00:14,982 --> 00:00:16,883
♪ HE'S ON HIS WAY
TO THE THRONE ♪
This is what my program spits out:
1
00:00:04,703 --> 00:00:07,007
HI, THERE. JUST WANTED
TO WELCOME YOU TO MY SHOW
2
00:00:07,073 --> 00:00:09,042
STARRING ME--KUZCO.
3
00:00:09,109 --> 00:00:11,845
SO, NO CHANGING
THE CHANNEL. UNDERSTAND?
4
00:00:11,911 --> 00:00:14,981
NO CHANGING.
OK. THEME MUSIC.
5
00:00:16,950 --> 00:00:18,551
♪ HE'S ON HIS WAY
TO THE THRONE ♪
I'm pretty sure mine is correct, some of the conversions Subtitle Edit produces don't actually make sense. It misses line breaks for some reason I can't fathom and seems to treat each frame as a frame... which is wrong.
Anyway, if anyone has m4vs they'd like to try my program on, you first have to extract the captions with mp4box [-nhml TrackID:Full] then drop the .media file onto the EXE. The srt should be forthcoming.
http://www.mediafire.com/download/ycvc0h8lhxu05kw/Break_Final_Cut.exe
It was written in c# and it needs to be completely rewritten, I need to add Flash On (although I doubt iTunes would ever use it) and try to figure out how ASS subtitles work so I can try retain the positionings (and maybe the colours as well).
Getting Subtitle Edit to do it for me would work too...
Nikse555
19th April 2015, 19:18
@JayJayH: If you use "DirectShow" video player in SE (default) and install DirectVobSub/VSFilter (use 64-bit if you have 64-bit os) then you should have a "preview" (it's not an intentional feature ;)
@minhjirachi: I'll do some testing on win 8.1 (my normal computer runs win 7). Also, what happens if you convert all images to 8-bit before importing into Scenarist?
@ndjamena: cool :) I do know SE has some issues regarding/extraction 608/mp4 . If you want to help improve SE perhaps you could share the source code? Also, you can check SE mp4 extraction + 608 decoding - e.g. start from here: https://github.com/SubtitleEdit/subtitleedit/blob/master/src/Logic/ContainerFormats/Mp4/Boxes/Stbl.cs
(I've not seen any nice documentation about either mp4 subtitle extraction or c608)
minhjirachi
20th April 2015, 03:21
@minhjirachi: I'll do some testing on win 8.1 (my normal computer runs win 7). Also, what happens if you convert all images to 8-bit before importing into Scenarist?
After using the BDSup2Sub to convert 8-bit images. The Scenarist can import them fluently.
ndjamena
20th April 2015, 10:38
For what it's worth (which isn't much) here's the source code:
http://www.mediafire.com/download/5gbqrafm7n728c3/608_Captions.zip
Bear in mind, originally I was just trying to split the track into files containing individual frames so I could have a look at them, then I tried translating each of the "commands" into strings I could output into a text file, then I though I'd try applying the commands to a buffer and finally I though I'd try converting the buffer into an srt. So it's just one thing built onto another, I hadn't really figured out how any of it worked until I'd got a few srts out of the way so I had to patch it up on the way and make it do things it wasn't designed for, and I'm pretty sure there's still some things I'm missing. Basically the code is junk, testament to my learning process but otherwise rather worthless and needs to be rewritten from scratch. (I've started redoing the buffers so I can swap them easier and use the second channel but haven't gotten very far yet, so there's some unused code in there too.)
It's not obvious from any of the documentation I've found, I assume they don't bother mentioning it because when used with it's original transmission medium (ie analogue NTSC) it should be fairly bloody obvious.
http://en.wikipedia.org/wiki/EIA-608
OK,
It uses a fixed bandwidth of 480 bit/s per line 21 field for a maximum of 32 characters per line per caption (maximum four captions) for a 30 frame broadcast.
480bps / 30 fps = 16 bits = 2 bytes
So basically, line 21 of each field contains 2 bytes worth of information, which means each 2 bytes in 608 captions, regardless of how it's stored or how it's being transmitted, increases the current timecode by 1/30 of a second (which is when the next even field begins) and every 608 "frame" (2 bytes) is a P-frame that inherits the state from the previous frame. (Captions update 30 times a second, to convert to srt you need to figure out where the actual "display" memory changes, capture it's current state and convert THAT to srt, rather than the commands themselves.)
The way MP4 stores the captions is misleading, there are no frames in 608 beyond the 2 bytes on each line 21, which leaves the mp4 frames themselves as having durations of [Number_Of_Bytes / 60] seconds. If the timecode of the current mp4 frame is later than the end of the last mp4 frame then that's the equivalent of the time in between the end of the last frame and the beginning of the current frame being filled with nulls. If the end of the last frame overlaps the beginning of the current frame then that's an authoring error.
[end of caption] SWAPS the displayed memory with the non-displayed memory. So basically, if I load the words "Rabbit Season!" into Non Displayed Memory [resume caption loading] then send [end of caption] the contents of the Non Displayed Memory and the Display Memory will be swapped and "Rabbit Season!" will be displayed on the screen. If I then load "Duck Season!" into Non Displayed memory and then send [end of caption] it will now display "Duck Season!" if I then wait a few seconds and send [end of caption] again, the buffers will be swapped and it will display "Rabbit Season!" again. I don't know if iTunes will ever use it like that, but they can if they want to.
The only way to move to a new line is to explicitly tell it to, if you've filled the last column in a row (column 32) and are told to display a new character, the standard is to replace the character in column 32 with the new character and keep doing that until you receive a command to move the cursor. If they send a [backspace] command and you're in column thirty two, the recommendation is to determine if you've just written to column 32, if you have delete the contents of column 32, otherwise delete the contents of column 31. (I keep forgetting that I've left a bug in my code in regards to that, my code sets the column number to -32 once column 32 has been written in to, I was supposed to add math.abs functions to all the array coordinates so they don't error out but haven't got round to it yet [FIXED].)
[resume direct captioning] writes the following commands directly to the display buffer. So if someone's screaming the same word over and over with increasing volume, you can pop up a caption saying "You!" the first time, then for the second, switch to Direct Captioning and send just a "!" to make it "You!!" and then send another "!" each time they yell, again I'm not sure iTunes would use it that way, but they can.
Most of my code was written with nothing but Wikipedia as a reference, I did find another document that filled in a few blanks, but it's a PDF and it wastes so much effort explaining the differences between how each generation of 608 decoders handle each command that it's hard to find anything useful in there. It doesn't help that they've denied copying permissions, so I can't copy/paste important parts or convert it to a Word Document to make finding things easier.
Anyway, that's most of it. The positioning data makes flawless conversion to srt impossible in every possible case, I can imagine situations where it would be almost impossible to figure out the order in which words are said without user intervention:
Hello |Hi,
How are you? |I'm fine thanks.
{Sorry, the forum is removing all the spaces separating the captions}
It would make perfect sense as a caption when you're watching the video, I could probably think of worse and more likely situations if I tried. If it was a comedy and they were deliberately playing with the captions maybe...
{If left to me my program will likely never get finished, my head's not the most stable construct on Earth and I'm surprised I got this far}
{I had more, but my head hurts so I'll stop now.}
-edit- [Replace 60fps with (60000/1001) fps if you like, or 30 fps with(30000/1001) fps]
Thunderbolt8
27th April 2015, 12:51
some more SHD removal fixes:
WOMAN: <i>Mr. Sportello?</i>|- Mm-hm.
gets changed to:
<i>- Mr. Sportello?</i>
but should get changed to: <i>Mr. Sportello?</i>
--> the WOMAN: speaker information part is not taken into consideration for SHD removal in combination with the hyphen from the speech part after the line break.@Thunderbolt8: thx, the "remove text for HI" issue should be fixed in latest beta: http://www.nikse.dk/SubtitleEdit.zip (portable versionhey, thanks. however, could you please also include the special "—" hyphen character in all those hyphen removal related rulesets as well? Current example is:
Uh— Well, I, uh—
gets changed to
— Well, I —
its correct for the last part in which the hyphen has to remain after the ", uh" removal (because the "I" is still real speech information), but is not taken into consideration as part of the 'no real speech' SHD stuff with which it should be removed together if thats the only thing remaining, as seen at the first part of that line.
ndjamena
27th April 2015, 14:50
Grrr, I discovered CCExtractor could extract captions from MP4, so I thought I'd give it a try. I figured PGS would be the best output format to keep the positionings, there's an output format called "spupng" which looked promising, it created and xml and a bunch of png images of the captions, but so far I haven't found anything that can even open it, much less convert it to pgs. So I thought I'd try simple srt output, I tested it on a file I'd already converted using my crappy little program... AND THEY GOT THE DAMN TIMECODES WRONG. I know it's wrong, because I can play the files synced with iTunes, mine lines up exactly, theirs doesn't. It's CCExtractor!!! WTF???
Unfortunately I've encountered my first 608 with Roll-Up captions, for one thing they're not what I thought they were, for another it makes mincemeat of my program. I looked at my code and got embarrassed enough to remove it from mediafire, I'll pretend it was never there. I've started rewriting it into a single neater 608 Captions class... I doubt I'll manage to finish it though.
RECOMMENDATION: Service providers shall calculate whether a Backspace is being issued
before or after display in Column 32. If a character or mid-row code has been placed in Column 32,
send a transparent space or Delete to End of Row command (DER) to erase the 32nd column. If
Column 31 is also to be erased, send the transparent space or DER first, then the Backspace.
In general purpose captioning, the Backspace command should not be used as long as TC1
decoders continue to be supported.
---------------
The FCC rules specify that "a Backspace received when the cursor is in Column 1 shall be ignored," but it
does not specify how Backspace should be applied when it is received following a character displayed in
Column 32. Since the rules say, however, that "Backspace shall move the cursor one column to the left,
erasing the character or Mid-Row Code occupying that location," and since there is no Column 33, many
manufacturers have concluded that a Backspace received either before or after displaying a character in
Column 32 shall move the cursor to Column 31 and erase the character there. This application is legal under
the rules, and, although a different method might have been preferable, all decoders shall implement
Backspace in this manner. When erasing Column 31, the decoder may also erase any displayable
character or other code in Column 32 (as is currently done in TC2).
OK, I read that wrong, it's the people who send the captions that are supposed to monitor what's in column 32, not the ones who decode it, we're supposed to delete both columns 31 and 32. I need to rewrite that. I still can't figure out if PACs take up a space, or where I'm supposed to store their text style info... If I press process a backspace after a MidRow, do I cancel the style, or just remove the space and not remove the style unless there's another backspace? I don't know.
Thunderbolt8
24th May 2015, 01:32
it seems like quotation marks screw up the SHD line detection a bit:
- Cover him!|EAMES: Down! Down now! ---> - Cover him!|- Down! Down now!
working as intended
- "My father doesn't want me to be him."|EAMES: Exactly. ---> - "My father doesn't want me to be him."|Exactly
the hyphen indicating a change of speaker after the line break is missing
izanami
1st June 2015, 09:42
i dont see my font in waveline . pls help me how can i do
http://postimg.org/image/557fqwv9b/
Nikse555
1st June 2015, 15:51
@Thunderbolt8: thx, this case should be fixed in next update.
@izanami: could you test latest beta: http://www.nikse.dk/SubtitleEdit.zip ?
(the text in the waveform will appear double... hopefully one of them will be correct)
@ndjamena: Sorry, I've not had too much time to look at it... will have more time in about 14 days I think
minhjirachi
1st June 2015, 16:42
@Thunderbolt8: thx, this case should be fixed in next update.
@izanami: could you test latest beta: http://www.nikse.dk/SubtitleEdit.zip ?
(the text in the waveform will appear double... hopefully one of them will be correct)
@ndjamena: Sorry, I've not had too much time to look at it... will have more time in about 14 days I think
I don't know why the bit depth of the exported bdn/xml files always 32-bit. Or I think that you should add the BDSup2Sub to Subtitle Edit, which process BDN/XML files really good.
izanami
2nd June 2015, 09:26
Thank You Nikse555 . My font is ok in SE 3.4.6 . I really appreciate for your help :)
ndjamena
11th June 2015, 15:01
In case you find something useful in it here is the other document I was using:
http://www.mediafire.com/view/0j4whlx7obo7ejf/EIA-CEA-608.unlocked.pdf
(I've unlocked it.)
I suppose it's possible iTunes is buggy and CCExtractor time-codes are correct, but the muxing mode is Final Cut Pro, which is a program owned and written by apple... someone with a Mac could check how FCP handles the time-codes, if it agrees with iTunes then that's pretty much the end of the story. I couldn't figure out where CCExtractor was getting it's time-codes from but 608 captions are line 21 captions for NTSC analogue broadcasts, theoretically if you play the M4Vs back on an analogue TV you should be able to write the captions back into line 21, which isn't possible with the CCExtractor time-codes because you'd have to start sending the information for the first caption before you even begin playback of the file to get it to display at the right time, and all the other time-codes are off too.
CCExtractor is open source... Apparently. Beyond that VLC has a 608 decoder if you can find it.
https://wiki.videolan.org/VLC_Source_code/
http://ccextractor.sourceforge.net/about-ccextractor.html
My test program has been downloaded 3 times now... I probably should have deleted it but I guess it kind of works if all you want from it is SRTs and you don't use it on files with roll up captions... And unless I missed a program it does seem to be the only way to get the correct time-codes short of owning a MAC... well, this looks like it will do the job but it costs $900:
http://www.drastic.tv/index.php?option=com_content&view=article&id=211&Itemid=303
There's probably something better out there but blow if I can find it.
That's about all I have to contribute. :(
Dean007
21st June 2015, 15:22
Hi. I have a question. Can I sync subtitles with a press of a button like Subtitle Workshop has (alt+m).
For example; I load subs and a video, mark from which subtitle I want to sync, play the video to that subtitle and simply press alt+m and subtitles that were marked are all sync with the video.
kalehrl
21st June 2015, 18:39
There is 'visual sync' option where you use 2 points for sync - at the beginning of the video and at the end for best results.
ndjamena
19th July 2015, 22:20
The developer of CCExtractor:
I'm probably not doing it correctly... didn't have too many itunes samples to begin with.
VLC won't play them properly, MPC-HC won't play them at all, neither CCExtractor nor Subtitle Edit will extract them properly, Handbrake doesn't notice they're there...
Does anything that's not an Apple product actually work with these things?
ndjamena
21st July 2015, 01:01
FFMPEG can see them and attempts to convert them to ASS when muxing to an MKV but ultimately the subtitle track comes out empty.
Music Fan
28th July 2015, 20:48
I have a strange problem with version 3.4.7 when I make OCR on french DVB-SUB ;
"J'ai gagné !" becomes "♪ ai gagné !"
Most of the J' are considered as the music symbol ♪
It's strange because in the "All fixes" part in OCR window, I see the correct spelling on the left and the bad correction on the right, whatever I check or not "fix OCR errors" (I use french dictionnary).:confused:
I guess it means that the correct spelling is detected but is wrongly corrected for some reasons, while no correction is needed in this case.:confused:
And if I don't choose french dictionnary and choose none, this problem disappear and the J' stay as is.
But I need it because if I let on none, I get other errors that I don't get when I choose french dictionnary.:o
Does it mean the french dictionnary is bugged ?
Something else : after OCR is done, I click on OK and I get this message : "do you want to discard changes made in current OCR session ?"
If I choose no, nothing happens, and if I chose yes, the OCR window closes and the text appears in the main window with the changes, as usual, thus I don't understand why I get this message :confused:
raymondjpg
29th July 2015, 02:24
Something else : after OCR is done, I click on OK and I get this message : "do you want to discard changes made in current OCR session ?"
If I choose no, nothing happens, and if I chose yes, the OCR window closes and the text appears in the main window with the changes, as usual, thus I don't understand why I get this message :confused:
Looks like a bug to me. I've gone back to v3.4.6, but I'll gladly persist with v3.4.7 if, as you say, clicking "yes" does not result in changes being discarded.
jpsdr
29th July 2015, 08:53
@Music Fan
As french user also, i want to know if this issue is specific to 3.4.7, or does it happen also with 3.4.6 ?
Nikse555
29th July 2015, 09:07
@Music Fan: thx for reporting the discard message after pressing OK - it's a bug - should be fixed on latest version on GitHub and also here: http://www.nikse.dk/SubtitleEdit.zip (beta, no installer)
Also, could you email me the sub that makes problems with music nodes?
Music Fan
29th July 2015, 23:08
Thanks for the fix, no message anymore.
But OCR on the sup I created yesterday (Blu-ray sup export from TS without OCR) is less well done than with 3.4.7 on some lines : ? is sometimes considered as 'I
edit : I can't reproduce this problem, no problem with the ? now :confused:
Look at this sup file (exported in Blu-ray sup from a TS file, no OCR was done, created with v3.4.7) ;
http://www31.zippyshare.com/v/p6pYpzXt/file.html
Same subtitles but exported this time with your last beta version (again without OCR) ;
http://www25.zippyshare.com/v/81vGbafc/file.html
The result of the OCR with this sup is exactly the same than when done from the original TS.
@ jpsdr : I don't know, I didn't try 3.4.6.
Music Fan
3rd August 2015, 19:10
Nikse555, did you find why J' become ♪ (with the sup I posted here) ?
Thunderbolt8
3rd August 2015, 20:53
could you please provide an example how the syntac structure has to look for this and how it is supposed to work out exactly? "Remove text for HI - "Remove if text contains" now allows multiple items separated by comma or semicolon" ?
Music Fan
4th August 2015, 12:59
3.4.8 is released ;
3.4.8 (2nd August 2015)
* NEW:
* Added support for Blu-ray TextST - thx Timo/ndjamena
* Added "Google it" to spell check dialog
* IMPROVED:
* Updated Chinese Simplified translation - thx Leon
* Updated Danish translation
* Updated Croatian OCR fix replace list - thx diomed & xylographe
* Remove text for HI - "Remove if text contains" now allows multiple items separated by comma or semicolon - thx Jesper
* Added fix for invalid time codes in Avid (bug in Avid) - thx Xenophon
* All Google urls now uses https
* FIXED:
* Fixed "Discard" message in OCR when pressing "OK" (regression from 3.4.7) - thx Music Fan
* Fixed Blu-ray sup export with frame rate 23.976 (regression from 3.4.7) - thx Arjan
* Fixed remebered value from Tools -> Adjust all times (regression from 3.4.7) - thx GH
* Fixed GT by using https - thx Sopor
* Fixed crash after using "Split" - thx Krystian
* Fix for large data inside "Sami" files - thx hhgyu
* Fix for font tags without quote/apos in "Advanced Substation Alpha" - thx hhgyu
* Now comboboxes from "Remove text for HI " should save/restore last used value - thx Jesper
* Fixed several issues with format "CIP" output - thx Victor
* + Many minor fixes from Ivandrofly and xylographe
I still have my J' problem but it's ok if I uncheck "music symbol" in the OCR window ! I don't remember if this option was already present in previous versions.
Anyway, ♪ can also be converted to J' with the "multiple replace" option.
Thunderbolt8
4th August 2015, 14:57
I'd like to request an option for blocking parts of a sentence from HI removal if you feel they belong together. e.g. I have 'oh' in the interjections list, but I dont want to remove it from "oh, my god" or "oh well", "oh boy" and such. currently, I still have to go through each subtitle line marked for HI removal manually and see if there is some line I need to untick and process manually. its quite tiring, especially when doing it for whole TV series like for example all 86 episode subtitle files from the sopranos. this improvement could save me a lot of time here.
von Suppé
24th August 2015, 07:26
Hi Nikse555,
Is there a possibility to give a color to the breakline tag <br /> and italic tags <i> and </i> in the list view?
Regards
von Suppé
speedyrazor
1st September 2015, 17:26
I am asking this here because I know what an excellent and powerful application this is and would love to use it for this task.
I am trying to convert a .itt file to .ssa and then change the height of the subtitles, top and bottom.
I have attached 2 files, Subtitles_da.itt and Subtitles_da.ssa (attached to this post in a zip file).
I am trying to alter the vertical position of the bottom and top subtitles so that neither comes into the blanking of the video file. So basically I need to make the bottom subtitle higher and the top subtitle lower. I don't think this is possible to do in .itt files, so my question is how do I do this in the Sub Station Alpha file?
I have also attached some screen grabs to show what I am trying to do (attached to this post in a zip file).
Let me know if there's any more info I can provide.
I appreciate any help you could offer.
Kind regards.
von Suppé
4th September 2015, 11:15
Is it allowed to ask questions about movie-subs from BD?
cheers
von Suppé
S_E_New
8th September 2015, 18:43
Hi! I want to know if is there any command line for export .srt to sub/idx with a .bat with these specifications.
http://i.imgur.com/6jBtan4.png?1
:thanks:
ukendt
10th September 2015, 10:27
Is it allowed to ask questions about movie-subs from BD?
cheers
von Suppé
Sure!
von Suppé
12th September 2015, 09:49
Sure!
Thank you, ukendt.
I ripped subs from the movie BD Interstellar (2014) and in SE OCR'ed them for editing reasons.
Now, this film has both 16:9 and (approx.) 2.40:1 footage.
Remuxing to mkv I leave video intact so PAR will stay 1:1, AR will stay 16:9 with most of the footage having black bars.
So, when exporting to SUP, I'd like to have the possibility to give some subtitles a different bottom offset.
I know this can be manually done with BDSup2Sub, but the workflow is rather unhandy and quite laboursome to me.
I'd like the idea of (in this case) selecting the concerning subs and give them another offset than the rest.
Any ideas?
Cheers
Thunderbolt8
5th October 2015, 20:40
lines which contain only numbers (and characters as - + / { ...) are considered as uppercase lines in removal of hearing impaired stuff with "remove line if UPPERCASE" ticked. I guess this is not meant to happen.
Xebika
7th October 2015, 05:58
3.4.10 is released ;
3.4.10 (6th October 2015)
* NEW:
* Audio visualizer waveform filled - thx jdpurcell
* FIXED:
* Fixed crash in "Visual sync" - thx aMvEL / Bolshevik
* Fixed "Fix common errors in selected lines" - thx ingo
* Fixed audio visualizer "Seek silence" - thx jdpurcell
* Fixed audio visualizer "Guess time codes" - thx jdpurcell
* Fixed possible startup crash with tiny or bad video file - thx ttvd94
* Fixed alpha in ASS styles - thx ravi
* Fixed line splitter regarding unicode 8242 char
3.4.9 (3rd October 2015)
* NEW:
* New subtitle formats (XIF xml, Jetsen, NCI Timed Roll Up Captions and more)
* Ukrainian translation - thx Maximaximum
* Option to play a sound when new network message arrives - thx InCogNiTo124
* IMPROVED:
* Updated Portuguese translation - thx moob
* Updated Korean translation - thx domddol
* Updated French language file - thx JM GBT/xylographe
* Updated Hungarian translation - thx Zityi
* Updated Dutch translation - thx xylographe
* Updated German translation - thx xylographe
* Updated Romanian translation - thx Mircea
* Updated Polish translation - thx admas
* Updated Croatian OCR fix replace list - thx diomed & xylographe
* Generating of spectrogram is now *many* times faster - thx jdpurcell
* SubtitleListView: Enable double buffering to eliminate flickering - thx jdpurcell
* Added "Count" to "Find" dialog - thx ivandrofly
* FFMPEG audio extraction will now prompt for audio file if more than one
* Format PAC now includes "Chinese simplified" - thx Man
* Merge lines with same text now ignores casing - thx Michel
* Now keeps blank lines inside SubRip texts
* Better Croatian/Serbian language detection - thx aaaxx/xylographe
* Sync tools now display info about what is applied (factor and -/+ adjustment)
* Plugins are now allowed to return another format than SubRip
* Better resizing of list view in "Bridge gaps in durations" - thx ivandrofly
* Audio visualizer: don't generate wav if source is already wav - thx MM
* Audio visualizer: zoom with control key + scroll wheel - thx jdpurcell
* "Fix common errors" only shows English "i to I" fix for English language - ivandrofly
* Allow up to 10 mb subtitles in batch convert - thx Kymophobia
* FIXED:
* Fixed tags accumulating texts in Sami format (regression from 3.4.8 refact) - thx domddol
* Show all paragraphs in audio visualizer when zoomed out - thx jdpurcell
* Fixed bug in "Set end and offset the rest" that degraded performance more and more - thx jdpurcell/Leon
* Fixed several issues with format "Cavena 890" - thx Victor
* Better handling of some zero width Unicode spaces - thx Krystian
* Spell check issue with multiple occurrences of same word in one subtitle - thx Krystian
* Rounding issue in formats with duration in output - thx Victor/Jamakmake
* Fixed possible crash in setting regarding VLC path - thx xylographe
* Fixed shortcut key in French replace dialog - thx Claude
* Go to first empty line now also focuses it - thx Jamakmake
* Clear overlap messages in main window after "New" - thx domddol
* Minor fix for split / timed text 1.0 - thx Krystian
* Issues with italic+bold tags in image export - thx marb99/aaaxx
* Possible crash in export to DOST
* Time codes in export to DOST - thx Christian
* Fixed crash when cleaning spectrogram temp images
* Audio visualizer: Fix crash when using mouse wheel without audio - thx jdpurcell
* Audio visualizer: Fix crash when using shortcuts without audio
* Audio visualizer: Fix issues with HH:MM:SS:FF time code format - thx jdpurcell
* Audio visualizer: Stable end time when using HH:MM:SS:FF time code format - thx ing
* Audio visualizer: Fix new selection disappearing if scrolled out of the left - thx jdpurcell
* Inline margin is now loaded/saved in SSA/ASS (can only be edited in source view)
* Changed LAV Filters link from Google Code to GitHub - thx suvjunmd
* Fixed bug in FAB export regarding center alignment - thx felagund
* Some fixes for "Remove text for HI" - thx Rasmus
* Fix for added words no "don't break after" list - thx ivandrofly
* Fixed bug in batch convert filter (2+ lines)
* + Many minor fixes from ivandrofly and xylographe
varekai
7th October 2015, 11:39
Hello!
First time user of SubtitleEdit 3.4.9
Just recieved a new Blu-ray disc Mad Max: Fury Road.
There are 2 subtitles I'd like to edit,
one is in english SDH where I want to remove the SDH,
and one which is in my opinion poorly translated to swedish.
I extracted the sups (00100.track_4609.sup and 00100.track_4617.sup) with tsMuxer and then imported it in SubtitleEdit.
It does its job and I can see the sup images showing the text but for some reason which I don't understand it won't generate any srt text?
File-->Import/OCR Blu-ray (.sup) subtitle file...-->00100.track_4609.sup--> Hit 'Start OCR' button... and nothing!?
I think I'm missing something fundamentally so any hints and tips would be much appreciated.
Regards
Edit: Just noted there's a new version, will update! Thanks!
Boulder
7th October 2015, 11:48
I've found out that the other OCR option works better than Tesseract so you might want to try that (it's the same you use with DVD subs). You probably need to tweak the two parameters it has, but it's quite easy once you get the hang of it.
varekai
7th October 2015, 12:06
OK, will try that, thanks!
Lucius Snow
10th October 2015, 12:11
Hello all,
I've got a little problem with an arabic subtitle that i want to export with Final Cut Pro + Image sequences. When there's a sentence in italic, with <i> and </i>, it displays properly in the video preview. However, when exporting, the dot, the comma etc. will get inverted (finishing at the right of the sentence instead of finishing at the left).
Any ideas?
Thank you.
mbcd
10th October 2015, 14:03
First of all:
Thanks for that really cool application !! :thanks:
While usage I found some things that made me a little "boring", because of batch conversation.
OCR:
Is it possible to do batch-OCR with "picture-recognition" ?
I mean not fully automatic batch, but selecting e.g. 20 subs, and they were loaded and saved automaticly.
Now you have to do those steps manualy.
- Load one Subtitle
- Start OCR
- Save one Subtitle
For mass-conversation much work. If they were loaded and saved (with automaticly started OCR by picture-recognition), each after each other, it would be very nice.
Export:
Is it possible to do a batch-export for files (load and export already existing subtitles)?
By now you have to load each subtitle by hand and export each subtitle by hand.
Best Regards
Thunderbolt8
10th October 2015, 22:24
numbers are still considered as uppercase characters in 3.4.10 and therefore will be removed if a line consists of numbers (+special characters) only when "remove line is uppercase" is ticked in SHD removal.
Thunderbolt8
21st October 2015, 17:40
for some reason this line gets treated incorrectly with SHD removal while other lines of the same type are not. Guess the hyphen at the end of the first displayed line throws it off.
WOMAN: Excuse me-|ALAN: Tim, I gotta call you back.
--> Excuse me-|ALAN: Tim, I gotta call you back.
but should be: - Excuse me-|- Tim, I gotta call you back.
jpsdr
23rd October 2015, 12:41
I have some issues with OCR Tesserac. Sometimes, i have the result which is something like "D: oi", when the picture displayed
I've also opened an issue on github here (https://github.com/SubtitleEdit/subtitleedit/issues/1394).
I can provide several files which produce the issue, PM for them if you want.
Version is 3.4.10.
jpsdr
24th October 2015, 09:13
Issue have been identified and fixed, nice work, thanks. Now, just have to wait for the next release.
varekai
3rd November 2015, 13:12
Can anyone identify this font?
Looks really nice and I would like to use it in a video project.
http://i.imgur.com/30cBd4q.jpg
http://i.imgur.com/267zC61.jpg
Edit:
Think I found a close match to the font.
This one will be perfect.
http://www.fontspring.com/fonts/fontsite/microsquare
Thunderbolt8
22nd November 2015, 16:13
some more SHD removal inconsistencies, some parts of lines are not removed correctly when italics are involved
("remove text between ♪ and ♪" is ticked and "remove text if it contains: ♪,♪♪" is unticked)
correct:
- ♪ Was a good friend of mine ♪|- All right! ==> All right!
incorrect:
<i>- ♪♪[Continues ]</i>|- It's pretty strong stuff. --> <i></i>|- It's pretty strong stuff.
this one here is debatable, I guess its working as intended with leaving the hyphen because it is needed as signal for speech. it just looks a bit strange just to have the hyphen within the italics and not the rest of the line, as it was already strange with having the hyphen + [Nick] speech information in italics and not the rest as well, before. if youd ask me, the italics could be removed in such a case, but as said, according to the rules its probably working as intended:
- The meal is ready. Let's go!|<i>- [Nick]</i> J. T. Lancer! ==> - The meal is ready. Let's go!|<i>-</i> J. T. Lancer!
in this case here, the italics tag needs to be closed after the first line, because the closing tag gets removed along with the content of the line after the break and the italics tag extends over that line break (this should not be a problem in case of subtitles which have a separate open&close tag for each part of a subtitle line, before and after the line break symbol):
<i>- Here it is.|- Whoa!</i> ==> <i>Here it is.
it would be nice if the "remove text between" tick box for ♪ and ♪ could be extended to also have ♪,♪♪ and ♪,♪♪ as in case of the "remove text if it contains:" box right next to it already has. right now, the line
♪ Trotting down the paddock|on a bright, sunny day ♪♪
does not get removed, because the open and closing music symbol is different. the line would get removed if the "remove text if it contains: ♪,♪♪" box was ticked, but then a line like
- ♪ Was a good friend of mine ♪|- All right!
would get removed entirely as well (instead of only the part before the line break). so in such cases, if a subtitle file consists of different lines like
- ♪ Was a good friend of mine ♪|- All right!
and
♪ Trotting down the paddock|on a bright, sunny day ♪♪
theres currently no option to have this handled automatically correctly (which is not to remove the "all right" part in case of the first line and remove the 2nd line entirely), youd have to tick/delete stuff manually. this could be avoided if the options to select from of the first box "remove text between" would be extended from only including ♪ to include ♪,♪♪ as well.
in such a case here, if SHD information gets removed and a hyphen (— or double --) follows afterwards, a syntactic indicator of ending or linking a sentence like full stop (.), exclamation mark (!) or comma (,) which directly preceeds that SHD information should be removed then.
Oh. Oh, yeah. Um — ==> Yeah. —
btw. Id like to request again a kind of exclusion list (case sensitive) for the "remove interjections" list you can edit manually. in cases of "oh, boy" or "oh, my" I wouldnt like to have the "oh" removed (just "my" doesnt really make sense) while I would like to see it removed in all other cases. same for stuff like "Er" or "er" as SHD information of speech (like errrr....), I would want this gone but then again not "ER" for emergency room. in the end, youd have to check the whole SHD removal list manually just to avoid such small unwanted cases.
Thunderbolt8
22nd November 2015, 18:22
also, please in the "compare subtitles" window, please change the colour of how same lines are marked in the right window, of that subtitle file you are comparing against. currently, that colour looks like sand, its too difficult to spot from the white background. please turn it into something easier to distinguish from white.
Lucius Snow
26th November 2015, 16:42
Hello all,
I've got a little problem with an arabic subtitle that i want to export with Final Cut Pro + Image sequences. When there's a sentence in italic, with <i> and </i>, it displays properly in the video preview. However, when exporting, the dot, the comma etc. will get inverted (finishing at the right of the sentence instead of finishing at the left).
Any ideas?
Thank you.
Nobody to fix this bug? (:
Nikse555
26th November 2015, 22:40
Yo :)
In SE 3.4.10 alpha color values from SSA/ASS styles are used in File -> Export -> image based format - easy to make transparent background box etc.
@Thunderbolt8: thx for all the info :)
I've tried to fix the "remove text for HI" issues here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.10/SubtitleEditBetaTest.zip (portable version without installer, beta)
@jpsdr: thx, some nice fixes :)
@mbcd: The above beta has some improvements to "Tools" -> "Batch convert...". OCR might work and BD Sup output is brand new.
@Lucius Snow: I'm not really sure, but try Edit -> Revert RTL start/end... - you can also try the "Simple rendering".
Thunderbolt8
27th November 2015, 18:31
thanks, I'll try it and report back something might be still not working as intended
Lucius Snow
10th December 2015, 20:15
Thank you. It works wih "Revert RTL start/ end".
Rudde
11th December 2015, 15:44
We need this in the Subtitle OCR world.
http://gizmodo.com/a-new-ai-system-passed-a-visual-turing-test-1747500554
mariner
15th December 2015, 07:54
Greetings Nike.
There appears to be an issue with add duration in the latest built. The duration of last subtitle always ends up with 3.099 sec.
Many thanks and best regards.
mbcd
25th December 2015, 11:19
Yo :)
@mbcd: The above beta has some improvements to "Tools" -> "Batch convert...". OCR might work and BD Sup output is brand new.
And I am very happy about that.
Thanks a lot for the new BETA, your program is very well formed, you got the right job and ... yes ... you love it ;)
I am mainly working with Bluray-Subs and I think I found out some little problems:
1.
On Bluray TEXT-Output you cant add a userstyle yet, you can choose it later by Style-Id, but you cant define your own.
Or are they definied at another place for reuseability?
2.
I am missing some feature to change position of those captions.
SRT doesnt support different positions, but I think that this is an "important" feature.
Cool would be, if position is read out by OCR / Capturing an directly reused.
Also creating different styles and applying them in the editor (on different captions at the same time) would be cool.
Often you have some seconds in the beginning of a film where you can read actors-name or something else.
It would be cool to get a function to shift the text verticaly for some captions, so that subtitletext and filmtext are not on the same position for that case.
Best regards and a happy new Year to you and your family.
Marco
StainlessS
13th January 2016, 14:10
Yo Nikse555,
Great improvements in SE, love it :)
One thing though, Loading MP4/AVC video, 720x436@25-2.35:1 1600Kbps shows weird video
with centre section greyscale and left and right parts weird color versions (as if UV starts 2/3 way across
clip and wraps around to left).
Was same in v2.4.8 which I had installed previously.
Not any problem though in use.
Nikse555
14th January 2016, 21:04
@mbcd: I was not sure the Blu-ray TextST was used... I think you can change position via the "region style", so I've added the possibility to duplicate a "region style" - in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.10/SubtitleEditBetaTest.zip (last chance to test the beta before the final 3.4.11)
@StainlessS: Don't know much about video decoding... SE just plays via either DirectShow (using e.g. LAV filters) or via VLC.
Boulder
15th January 2016, 09:43
(last chance to test the beta before the final 3.4.11)
There's at least two reported issues that are not fixed yet in the beta. Please see the example subtitle file, on lines 401 and 439 there are two similar problems (words Yourpackage --> Your package and oftwo --> of two). Line 401 remains corrected but line 439 isn't corrected even if I change it manually.
https://drive.google.com/file/d/0BzeF_1syecQwYmpkRmdoQnJUM3M/view?usp=sharing
The other one is that if you import NTSC subtitles via IFO, the timestamps are incorrect. If you first use VSRip to get the IDX and SUB file and import that, you get correct timestamps. PAL subs seem to be fine imported via IFO.
Ghitulescu
15th January 2016, 10:04
I am missing some feature to change position of those captions.
SRT doesnt support different positions, but I think that this is an "important" feature.
Cool would be, if position is read out by OCR / Capturing an directly reused.
Also creating different styles and applying them in the editor (on different captions at the same time) would be cool.
Often you have some seconds in the beginning of a film where you can read actors-name or something else.
It would be cool to get a function to shift the text verticaly for some captions, so that subtitletext and filmtext are not on the same position for that case.
If the SRT does not support positioning (other formats do) how do you propose this to be implemented?
I disregard the usability of text subtitles in BD world. It may only work 100% correct if only the 26 English letters and numbers are used. All others depend on the settings and features of the player.
I only miss a possibility to change the colours for PGS subtitles. This I consider a basic function.
Nikse555
16th January 2016, 12:13
SE cannot edit image based formats - only import and export ( and only limited styling support in export - http://www.nikse.dk/SubtitleEdit/Help#export )
Also, SE 3.4.11 is out: https://github.com/SubtitleEdit/subtitleedit/releases
Shortened change log:
3.4.11 (16th January 2016)
* NEW:
* New "Binary Image Compare" OCR
* Can now import TextST from Matroska (.mkv) files - thx Rudde
* Can now import "Timed Text image" files - thx Mouna
* Added "Blu-ray sup" output in "Batch convert" - thx Dan
* IMPROVED:
* Some improvements for default ASS/SSA style - thx Michael
* Batch convert now remembers ASS/SSA style - thx Thunderbolt8
* TextST: Can now duplicate a "Region style"
* OCR Window: "Unknown words" is now tab one in the "log tab control"
* FIXED:
* "Normal (remove formatting)" don't remove Unicode chars - thx Fotis
* Fixed bug in spell check "word replace" - thx Sami
* Fixed post OCR correction regarding dashes in dialog - thx Rudde
* Fixed bug in OCR regarding scrambled line - thx jpsdr
* Several fixes in remove text for HI - thx Thunderbolt8
* Some fixes for the "Advanced Sub Station Alpha" properties window
* Fixed bug in color chooser - thx Thunderbolt8
* Fixed bug in "Remove interjections" - thx Leinad4Mind
* Fixed some crashes in OCR window - xylographe/heforfree
The "Binary Image Compare OCR" works better with some subtitles than Tesseract (also, Tesseract has a problem with "...").
Tip: If a letter is recognized wrong, then right click in the list view and choose "Inspect compare matches for current image" and "add better match" or update/delete the old match.
Ghitulescu
16th January 2016, 13:28
SE cannot edit image based formats - only import and export
PGS and TXT subtitles have different underlying basic requirements.
If these cannot be fulfilled what's then the purpose of editing them?
I am not criticizing, just wondering...
Boulder
16th January 2016, 17:56
There's at least two reported issues that are not fixed yet in the beta. Please see the example subtitle file, on lines 401 and 439 there are two similar problems (words Yourpackage --> Your package and oftwo --> of two). Line 401 remains corrected but line 439 isn't corrected even if I change it manually.
https://drive.google.com/file/d/0Bze...ew?usp=sharing
The other one is that if you import NTSC subtitles via IFO, the timestamps are incorrect. If you first use VSRip to get the IDX and SUB file and import that, you get correct timestamps. PAL subs seem to be fine imported via IFO.I've tested these ones, and they still occur. The first one is weird, because the correction is exactly the same for both but it only works for the first item :confused:
Nikse555
16th January 2016, 19:50
@Boulder: I do think I added those fixes to the "OCR fix list" (see Options -> Settings -> Word lists). Perhaps the installer don't overwrite the "eng_OCRFixReplaceList.xml" file as it should...
If you add a lot of OCR fixes do email me the file "eng_OCRFixReplaceList_User.xml" :)
Boulder
16th January 2016, 19:57
I didn't overwrite the file just because I have added stuff as I've OCR'd things :) I'll do a compare to see what the changes are and send you my file.
Anyway, the latter correction is not kept if I do it manually while doing OCR; that is, change "oftwo" to "of two" and select Change.
Nikse555
16th January 2016, 20:08
@Boulder: Later version of SE writes the user additions to the user file, so it should be simpler to upgrade.
Only the "Change all" button (+ "use always") adds the correction to the OCR fix list - http://www.nikse.dk/SubtitleEdit/Help#importvobsub
Boulder
16th January 2016, 20:12
Not working for me, I just tested with that troublesome subtitle. If I choose Change all or Use always at that point, the item doesn't appear in the user xml file (and is not corrected even though the status bar says so). It seems that the italics are the problem, or the first word of a sentence. I've never seen this issue with words in the middle of a sentence.
EDIT: I use the portable EN (or FI) version.
Nikse555
17th January 2016, 09:30
Not working for me, I just tested with that troublesome subtitle. If I choose Change all or Use always at that point, the item doesn't appear in the user xml file (and is not corrected even though the status bar says so). It seems that the italics are the problem, or the first word of a sentence. I've never seen this issue with words in the middle of a sentence.
Hm, sounds like a bug but I cannot re-create this...
http://www.nikse.dk/SE1.png
Boulder
17th January 2016, 10:34
I did have SE installed at some point, could that have messed things such as paths etc. up? I have uninstalled it and also tested a fresh portable installation, but it's still a no-go.
Did you try loading srt file I linked to and OCRing that one without having those automatic fixes in the corrections list?
LouieChuckyMerry
18th January 2016, 02:34
Hello and many thanks for Subtitle Edit :) . I just updated from 3.4.10 to 3.4.11 on my two laptops (Win 7 x64) and noticed that the the SE icon for .srt files has changed to the default icon for .txt files (Windows Notepad icon) on both setups. But only for .srt files; every other format I've set to default open with Subtitle Edit still have the SE icon (.ass, .idx, .sub, .sup, .ssa). I wondered if anybody else has experienced this or if my laptops are finally plotting against me after years of abuse ;) . Thanks.
Nikse555
20th January 2016, 22:53
@Boulder: I cannot ocr an srt file... or did you mean Fix common errors - fix common ocr errors? (I tried to export the srt to bluray sup and ocr that one)
@LouieCheckyMerry: did you use the installer version and did you check the new 'associate srt files with SE' checkbox?
LouieChuckyMerry
21st January 2016, 04:25
@LouieCheckyMerry: did you use the installer version and did you check the new 'associate srt files with SE' checkbox?
Yes, I used the installer version and ticked the new associate box. I also tried resetting .srt files to default open with Subtitle Edit but they were still set to default open with Subtitle Edit so nothing changed. Eventually I saved the Subtitle Edit AppData folders on both laptops then uninstalled-reinstalled Subtitle Edit and all is OK. I noticed, however, that .srt and .sup files now have a "newer" looking icon while .ass, .ssa, .sub, and .idx files have the old icon. Is this by design? I don't really care about the subtly--;)--different icon for .srt and .sup files, I'm simply curious if this is as it should be :) .
On a related note, a couple-three releases ago (I'm a bit slow, ha ha) the "Fix common OCR errors (using OCR replace list)" started changing (without the quotes) "Mr." to "Mr...". I've added "Mr...-->Mr." to "Settings/Word list/OCR fix list" but that didn't change anything; the only work around I can come up with is to use "Replace" with "Mr..." and "Mr." then hit "Replace all", which of course means that I then have to repeat "Fix common errors...", sort by "Function", deselect all the incorrect "Mr." to "Mr..." lines, then apply the affected "Break long line" or "Merge short line (all except dialogs)" as necessary. Certainly not the end of the world but quite annoying. Any ideas how to fix this? Thanks again.
Boulder
21st January 2016, 04:52
@Boulder: I cannot ocr an srt file... or did you mean Fix common errors - fix common ocr errors? (I tried to export the srt to bluray sup and ocr that one)Sorry, didn't mean OCR but spell checking on the srt file.
mariner
21st January 2016, 12:19
Greetings Nik.
There appears to be an issue with add duration in the latest built. The duration of last subtitle always ends up with 3.099 sec.
Many thanks and best regards.
Greetings Nik. Many thank for the new build
1. Did you have a chance to look at the above issue?
2. Does SE support command line mode to convert text srt to blu-ray sup, with all the formatting options available in the GUI?
Many thanks and best regards.
S_E_New
24th January 2016, 10:33
Cool would be, if position is read out by OCR / Capturing an directly reused.
It would be nice if it could be with sub/idx and sup.
And also Add "VobSub sub/idx" output in "Batch convert" :)
:thanks:
mbcd
25th January 2016, 11:11
Thanks for those Updates Nikse555 :thanks:
Meanwhile I figured out that there is an extended SRT-Standard which includes textposition.
Should be something like this:
2041
01:31:56,887 --> 01:31:58,012
{\pos(250,270)}<i>Come on. Who needs</i>
<i>something like this?</i>
2042
01:31:58,096 --> 01:32:01,223
{\pos(250,270)}<i>Who's the target demo</i>
<i>for a drug like Sustengo?</i>
Also BDN-XML has this option as you know, but I didnt figured out how to take in text there.
My personal goal was to scale subtitles up and down, but imagescaling look too ugly, so I thought to make OCR to do it in an elegant way.
Bluray-Text-Output is not a must-be, it is a nice feature, but not used because of some problemes. I thought it solves my Problems so I tried to use it, but also if it dous not result back in an .sup it does not make sence.
I want to take both in one file, but atm this is not possible. With SRT I loose position, with BDNXML I loose text ... but away from that, the textposition is generally "thrown away" after OCR ... so this is the real problem.
johner23
8th February 2016, 02:18
Hi, dear all!
@ Nikse555
Subtitle Edit can open, create, process or edit .stl files? If not, can you add such features on its engine?
See the thread below to get more information about the subject:
---> http://forum.videohelp.com/threads/288632-How-to-open-and-edit-stl-file
Best regards. Thanks for your time!
devil (johner)
arslan
14th February 2016, 12:40
Hello everybody, this is my first post.
Thank you Nikse555 for the excellent program!
I would like to suggest a feature that I believe could be useful. Would it be possible to implement saving and loading of user-defined settings as profiles?
Here is what I mean: if one produces subtitles for a TV show, it is 25 fps, 32 characters max, 160 milliseconds minimum pause, etc. When one produces subtitles for theatre movies, it is 24 fps, 40 characters, etc. So, instead of manually adjusting the settings every time, would it be possible to save a complete settings profile for TV, another one for movies, and for any other purpose? Then one would just load a defined profile and be ready for work.
Cheers!
mariner
14th February 2016, 14:02
Greetings Nik.
Appreciate if you could kindly look into a couple of issues when importing subtitle in xml format:
1. Time codes appear to be messed up.
2. When converting 2 decimal ms to 3 decimal places, leading zero is padded instead of trailing zero.
ie, .xx becomes .0xx instead of .xx0
Many thanks and best regards.
doomer9er
19th February 2016, 00:55
Hi, does anyone have a good set of Tesseract data & Dictionary files to help OCR .idx/.sub files into srt? The standard set that comes with Subtitle Edit is poor and has a lot of trouble with italics. Please help. Thanks.
S_E_New
20th February 2016, 07:51
It would be nice if it could be with sub/idx and sup.
And also Add "VobSub sub/idx" output in "Batch convert" :)
:thanks:
Hello everybody, this is my first post.
Thank you Nikse555 for the excellent program!
I would like to suggest a feature that I believe could be useful. Would it be possible to implement saving and loading of user-defined settings as profiles?
Here is what I mean: if one produces subtitles for a TV show, it is 25 fps, 32 characters max, 160 milliseconds minimum pause, etc. When one produces subtitles for theatre movies, it is 24 fps, 40 characters, etc. So, instead of manually adjusting the settings every time, would it be possible to save a complete settings profile for TV, another one for movies, and for any other purpose? Then one would just load a defined profile and be ready for work.
Cheers!
And if you could add others featuretes for VobSub sub/idx like:
-Auto-wrap lines
-Shadow in Imagen Settings
-The settings, can be stored in the Profile menu.
-Each subtitle's font, color, size and position can be set independently.
-And the possibility to insert pictures.
:thanks:
Boulder
21st February 2016, 17:46
@Boulder: I cannot ocr an srt file... or did you mean Fix common errors - fix common ocr errors? (I tried to export the srt to bluray sup and ocr that one)
Sorry, didn't mean OCR but spell checking on the srt file.Were you able to confirm the problem? I still see it from time to time when OCR'ing subs..
Also, where is the "Use always" or "Change all" word pair saved? I have used it extensively but I don't see the pairs being added to any xml file or in Settings --> Word lists --> OCR fix list.
Betsy25
28th February 2016, 00:52
If people are willing to give the Testbuild with mpv player (x64 bit version ONLY for now!) (https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.11/SubtitleEditBeta.zip) a try, and if problems occurs or issues or recommendations, don't hesitate to post in the related issue tracker. (https://github.com/SubtitleEdit/subtitleedit/issues/1579) :o
Lucius Snow
1st March 2016, 19:49
Hello Nikse555,
Can you please give more details about this improved feature coming for the next release?
Export to FCP+image now allows full frame + bottom border - thx Mathieu
Full frame is native resolution PNG file (i.e. 1920x1080) ? Which is not the case at the moment.
Thank you.
Thunderbolt8
1st March 2016, 23:08
does anyone know if the subtitle format of UHD blu-rays will be the same as that of normal BDs?
LouieChuckyMerry
12th March 2016, 12:14
Happy Saturday! I often use Subtitle Edit in conjunction with MPC-BE and a programmable mouse to, well, edit subtitles, taking advantage of MPC-BE's "Auto-reload modified subtitles" feature. Some change in Subtitle Edit from version 3.4.10 to 3.4.11 has made this impossible. With 3.4.10 I could save the changes to the subtitle and it would immediately appear in MPC-BE; however, with 3.4.11 when I save the changes in a subtitle the subtitles disappear in MPC-BE and I need to manually reload the subtitle. Did Subtitle Edit change the way it informs Windows 7 that a file has been saved? I apologize for my ignorance, thanks for any help.
Nikse555
12th March 2016, 23:28
@Lucius Snow: When using File -> Export -> FCP+image... you can now check "full frame" which will make the subtitle images full frame (full screen size for the chosen resolution - you can test it in the beta link below)
@LouieChuckyMerry: yeah, that's a bug in 3.4.11 which should be fixed in the soon to be released 3.4.12... feel free to test the latest portable beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.11/SubtitleEditBeta.zip
LouieChuckyMerry
13th March 2016, 02:24
Thank you! :thanks:
Nikse555
20th March 2016, 10:41
3.4.12 is out - download/change log: https://github.com/SubtitleEdit/subtitleedit/releases
New subtitle formats added (must be getting close to 250 formats) + added support for reading "S_DVBSUB" from mkv files + fixed many minor bugs + added support for mpv video player.
sheppaul
21st March 2016, 02:38
Subtitle edit seems not respect the menu/dialog font settings from system. It looks too small and is really hard to read in my UHD monitor. Is the font size fixed to 8 point? I'm currently using 10 point with 100% dpi in Windows 10.
Lucius Snow
28th March 2016, 13:33
@Lucius Snow: When using File -> Export -> FCP+image... you can now check "full frame" which will make the subtitle images full frame (full screen size for the chosen resolution - you can test it in the beta link below)
Great thing! Thank you.
hello_hello
28th March 2016, 16:16
Hi. Thanks for Subtitle Edit.
Is there a way to make the dictionary selection for spellcheck "stick" between spellchecks? For example I generally use the English U.K. dictionary but the default is English U.S. If I initiate a spellcheck I assume it commences using the English U.S. dictionary but if no spelling errors are found there's no opportunity to change dictionaries. To work around that I've been deliberately adding a word with incorrect spelling to the first line of a subtitle, but is there a way to set the default spellcheck dictionary I'm too silly to find?
Thank you.
S_E_New
30th March 2016, 23:38
Just to let you know:
• VobSub via command line is not working.
I try via command line SubtitleEdit.exe /convert "*.srt" VobSub and this is the result:
target format 'VobSub' not found!
0 file(s) converted
I also try with SUP SubtitleEdit.exe /convert "*.srt" Blu-raysup and it worked.
• Batch converter VobSub: Border style
The first line of subtitles is very thin although the configuration is Normal, width=6. This only happens with batch converter and not in Export>VobSub.
mariner
1st April 2016, 07:15
Greetings Nikse555. Thanks for the new update.
1. Adjust Durations still not working for the last subtitle entry.
2. BD SUP time code seems to be ~40ms off.
Thanks and best regards.
Nikse555
1st April 2016, 15:46
@hello_hello: OK, I've tried to fix it in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.12/SubtitleEditBeta.zip - better?
@S_E_New: "SubtitleEdit.exe /convert *.srt vobsub" works fine here... the border was a bug (in first line) which should be fixed in above beta.
@mariner: "Adjust durations" still works fine here for last entry... how can I re-produce the error?
mariner
1st April 2016, 17:05
@mariner: "Adjust durations" still works fine here for last entry... how can I re-produce the error?
Greetings Nikse555.
The increment is capped at .999 for the last entry.
Nikse555
1st April 2016, 18:50
Greetings Nikse555.
The increment is capped at .999 for the last entry.
OK, saw it now, thx :)
Duration adjustment in "Seconds" with values higher than one second will trigger the bug - should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.12/SubtitleEditBeta.zip
von Suppé
13th April 2016, 19:40
Hi Nikse555,
The plugin for removing the hyphen in the first dialog line works when there are 2 dialog lines, both with hyphens.
It would be nice to be able to remove the hyphen in the first line anyhow - so even if there's only one line.
Any chance for this?
Cheers
von Suppé
Thunderbolt8
15th April 2016, 23:14
this should already work with fix common error tab
von Suppé
16th April 2016, 10:53
this should already work with fix common error tab
Jeez, I feel so stupid...
Must've been braindead or something not noticing this, using this software for years.
Thanks & sorry :o
Edit: I tried this "Fix (remove dash) lines beginning with dash (-)" in Tools. When applied, it also removes the dash in the second one, for example:
Oh, no.
- Let it go.
becomes
Oh, no.
Let it go.
This is not what I want; I want the dash removed in only the first line, no matter how many lines there are or if hyphens are used in the second.
Any thoughts?
Thunderbolt8
18th April 2016, 12:24
you have to untick these cases manually then
Thunderbolt8
24th April 2016, 02:01
would it be possible to add more precise .srt italic brackets adding when importing a sub/idx dvd subtitle file? italic brackets only seem to be added for an entire line, but not parts of a line as sometimes is supposed to be.
hello_hello
25th April 2016, 01:20
To work around that I've been deliberately adding a word with incorrect spelling to the first line of a subtitle, but is there a way to set the default spellcheck dictionary I'm too silly to find?
@hello_hello: OK, I've tried to fix it in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.12/SubtitleEditBeta.zip - better?
Sorry about the very slow reply, but yes, that seems to be much better. :)
Thanks.
von Suppé
25th April 2016, 08:18
would it be possible to add more precise .srt italic brackets adding when importing a sub/idx dvd subtitle file? italic brackets only seem to be added for an entire line, but not parts of a line as sometimes is supposed to be.That is exactly why I earlier asked to make it possible to hi-light / color several tags, so you can track & check them manually easier.
Music Fan
25th April 2016, 09:23
Is there a way to add a black box around subtitles ?
mood
25th April 2016, 22:33
Is there a way to add a black box around subtitles ?
only in export window, select the text and right click on mouse and choose "box-single line" or "box - multi line" or in border style choose box
Music Fan
26th April 2016, 18:09
Thanks, well hidden. Is there a way to apply it to all lines ?
edit : yes, choosing "one box" in border style, as you said (I didn't understand directly that its effect was different than "box-single line" or "box - multi line" in the thext window).
mood
26th April 2016, 18:12
Thanks, well hidden. Is there a way to apply it to all lines ?
yes, in export window border style choose box is applied to all lines
edit: the difference is "single line" separates the break line, "multi line" uses black box with break line
is the same thing with space in between line with and without break line
It was better if "single line" it was only for the line without break
Music Fan
26th April 2016, 18:22
Yes, that's great (I found just before seeing your answer but thanks anyway).
zioneed
2nd May 2016, 17:03
Question on PGS subs and flag "forced".
Hi all, wondering if anyone can share some light on this:
SE is terrific with reading and converting PGS to srt.
I am in process of re-encoding some Game of Thrones bluray discs and I have found something weird: PGS track only for subs (ita, standard subtitles set, whole episode) but when I watch the video on PC using mpc-hc the subtitles menu offers the chance to pick up ITA-forced.
This is really puzzling me: no matter if I use dvdfab or staxrip to rip but, granted, there is just ONE ita subs choice.
Selecting "ita forced" from player I can actually have the proper subtiles only when the language is the one they use as native language (dothraki or whatever). A few lines only.
Question is: how can I select these only from the whole file? There is any "flag" or anything I can work on SE? any tick or something?
Really bizarre situation, cant work it out and I'm not a noob at these things.
Many thanks in advance!!!
sneaker_ger
2nd May 2016, 20:42
In a single PGS track you can have lines flagges as forced while others are not. MPC-HC's LAV Splitter can present these as a virtual track to select. Since it does not actually exist as a separate track you won't see it in other programs. You can check that by (during playback in MPC-HC) going into "Play"->"Filters"->"LAV Splitter [...]" and toggling "Enable Automatic Forced Subtitle Stream". Restart MPC-HC to test new setting. If the track disappears when option is turned off you know it was only a virtual track.
For such PGS subtitles in Subtitle Edit you only have to tick "Show only forced subtitles" at the bottom of the PGS import/OCR window.
zioneed
3rd May 2016, 05:48
In a single PGS track you can have lines flagges as forced while others are not. MPC-HC's LAV Splitter can present these as a virtual track to select. Since it does not actually exist as a separate track you won't see it in other programs. You can check that by (during playback in MPC-HC) going into "Play"->"Filters"->"LAV Splitter [...]" and toggling "Enable Automatic Forced Subtitle Stream". Restart MPC-HC to test new setting. If the track disappears when option is turned off you know it was only a virtual track.
For such PGS subtitles in Subtitle Edit you only have to tick "Show only forced subtitles" at the bottom of the PGS import/OCR window.
Thanks, it works fine!
guess it's too late but what I do is convert sup files to srt. Once done it's too late? If I have the srt file with thw whole set of lines there's any way to pick up forced only or the flags are gone?
I guess they're gone..
sneaker_ger
3rd May 2016, 10:23
They are gone.
Nikse555
12th June 2016, 16:43
Subtitle Edit 3.4.13 is out just now: https://github.com/SubtitleEdit/subtitleedit/releases
* Some minor improvements for OCR via "Binary image compare"
* Fixed end-time for (some) TS files - thx kurosu
* Added support for spu/png OCR - thx claunia
* Drag'n'drop to subtitle list now supports ifo/vob/mp4/mks (with subs) - thx Budman
* Fix updating of main list view after "Remove text for HI" + "Apply" - thx Henry
* Fix for border size in vobsub export via cmd line - thx S_E_New
* Adjust duration fixed for last subtitle - thx mariner
* Read more variations of "S_DVBSUB" from mkv files - thx dreammday
* Fixed some "Remove text for HI" issues - thx Thunderbolt8
* Fixed crash in OCR when changing "Use color" - thx kurosu
von Suppé
2nd July 2016, 12:12
Hi Nikse
I'm playing around and testing .ass subtiltes to export as SUP.
Now I'm using two different "ASSA" styles having different fonts. In the SUP export window I put different Line height settings for each style. SE doesn't appear to do honour that.
Example: I set line 1 with style 1 to Lineheight 62. After that I set Lineheigt to eg. 73 for line 2 with style 2.
When I go back to line 1, its Lineheight is changed back to a default 101, also when I put another number for style 2.
Is this something that can be fixed? I'd love to be able to lock a dedicated Lineheight to an ASSA style.
cheers
Ennio
mbcd
10th July 2016, 15:26
Yes, I tried that too, with the styles of .ass there are some little problems. Color is taken, but position (different borders in .ass-styles) are not recognized by output.
So to say it in short. The styles of .ass are not taken correctly by output to .sup.
mbcd
16th July 2016, 14:29
There is also a "BUG" with the waveforms.
If you use VLC as Mediaplayer, and you have a source which contains more than one audio-track, you get still shown the first audio as waveform.
That is not good, because you cant adjust and time subtitles to another audio-language, you still need the waveform of the actual choosen audio-track.
Nikse555
16th July 2016, 22:19
@von Suppé: Could you test latest beta - it should remember line height for each style: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.13/SubtitleEditBeta.zip
@mbcd: The above beta should also support pos(x,y) tags (as well as borders/italic/font size/font face/bold/color). If your video contains more than one audio track SE should prompt for audio track... be sure that you have latest VLC or FFMPEG.
von Suppé
17th July 2016, 10:57
Awesome. I go testing.
Much obliged for your work, Nikse :-)
van Suppé
Thunderbolt8
17th July 2016, 14:21
would it be possible to modify batch conversion to do batch processing without conversion as well? or just alternatively add a batch processing option. the point is that when using batch conversion as a kind of batch processing just for applying stuff like 'removing text for HI' or 'fix common error settings' for multiple files then batch format conversion is applied automatically as well. this is problematic in case of .ass subs which have been created with dvdsubextractor with the keeping the original line positions on screen option when using the .ass output format. in that case the header is different from the default one I use for all other standard .ass files and musnt be changed. otherwise the original line positions cant be shown. in such a case Id have to copy & paste the default header I use in subtitleedit elsewhere, then copy&paste the header of such a .ass file, then do the conversion/processing and afterwards restore the default header information again. this is all a bit clumsy and could be avoided by simply having the option to do batch processing without being forced to use conversion as well.
Nikse555
17th July 2016, 16:29
@Thunderbolt8: thx, was a bug... should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.13/SubtitleEditBeta.zip
mbcd
17th July 2016, 20:37
@von Suppé: Could you test latest beta - it should remember line height for each style: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.13/SubtitleEditBeta.zip
@mbcd: The above beta should also support pos(x,y) tags (as well as borders/italic/font size/font face/bold/color). If your video contains more than one audio track SE should prompt for audio track... be sure that you have latest VLC or FFMPEG.
Will test it out if output is taken to SUP !!!
BTW: Could it be that you are using some wrong EPOCH in creating SUP? I created several bluray-subs, and eac3to reports every time that there are only 1 caption inside, but there are always hundreds. Extractred SUP from blurays report the correct value (higher values) with eac3to. I think there is some fault.
EDIT: Looks like a caption of SubtitleEdit is never "closed" correctly in SUP.
amayra
23rd July 2016, 11:27
thanks for your hard work and i have question do you update beta ver frequently ?
PS: i like how your repository always active
Nikse555
23rd July 2016, 15:23
EDIT: Looks like a caption of SubtitleEdit is never "closed" correctly in SUP.
Very possible as I've never seen the specs and SE sup code is loosely based on bdsuptosub. Do you know how a cap should be closed correctly?
@amayra: The beta is normally updated 2-3 three times a week. Thx, always much to do :)
von Suppé
3rd August 2016, 19:54
@von Suppé: Could you test latest beta - it should remember line height for each style: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.4.13/SubtitleEditBeta.zip
During testing I encountered same problem again. Putting different line-height numbers seems to be remembered initially, but after scrolling upwards manually the numbers go back to a default 101 again.
Also, when putting the bottom offset in ass subtitle styles, both output-SUP and preview in SE's output-window are too high.
Cheers
von Suppé
edcrfv94
22nd August 2016, 22:53
SubtitleEdit can batch extract DVB-SUB from a ts file and convert to a sup file?
Music Fan
4th September 2016, 16:11
Hi,
I'd like to know if there is a way to remove completely the double lines in subtitles for hearing impaired.
Here is the debut of the original subtitle (extracted with CCExtractor) containing a lot of lines, nearly one for each new word ;
1
00:00:01,920 --> 00:00:02,120
Heureux de vous retrouver pour ce
2e numéro de
2
00:00:02,160 --> 00:00:02,360
Heureux de vous retrouver pour ce
2e numéro de la
3
00:00:02,400 --> 00:00:02,560
Heureux de vous retrouver pour ce
2e numéro de la 11e
4
00:00:02,600 --> 00:00:02,800
Heureux de vous retrouver pour ce
2e numéro de la 11e saison
5
00:00:02,840 --> 00:00:03,000
Heureux de vous retrouver pour ce
2e numéro de la 11e saison de
6
00:00:03,040 --> 00:00:03,240
Heureux de vous retrouver pour ce
2e numéro de la 11e saison de "On
7
00:00:03,280 --> 00:00:03,440
2e numéro de la 11e saison de "On
n'est
8
00:00:03,480 --> 00:00:03,680
2e numéro de la 11e saison de "On
n'est pas
9
00:00:03,720 --> 00:00:03,880
2e numéro de la 11e saison de "On
n'est pas couché"
10
00:00:03,920 --> 00:00:04,120
2e numéro de la 11e saison de "On
n'est pas couché" avec
11
00:00:04,160 --> 00:00:04,320
2e numéro de la 11e saison de "On
n'est pas couché" avec notre
12
00:00:04,360 --> 00:00:04,600
n'est pas couché" avec notre
nouveau
13
00:00:04,640 --> 00:00:04,800
n'est pas couché" avec notre
nouveau duo
14
00:00:04,840 --> 00:00:05,040
n'est pas couché" avec notre
nouveau duo :
15
00:00:05,080 --> 00:00:05,320
n'est pas couché" avec notre
nouveau duo : Vanessa
16
00:00:05,360 --> 00:00:06,120
n'est pas couché" avec notre
nouveau duo : Vanessa Burggraf
17
00:00:06,160 --> 00:00:06,360
n'est pas couché" avec notre
nouveau duo : Vanessa Burggraf et
18
00:00:06,400 --> 00:00:06,560
nouveau duo : Vanessa Burggraf et
Yann
19
00:00:06,600 --> 00:00:08,560
nouveau duo : Vanessa Burggraf et
Yann Moix.
20
00:00:08,600 --> 00:00:08,880
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir
21
00:00:08,920 --> 00:00:09,120
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir à
22
00:00:09,160 --> 00:00:12,640
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir à tous
23
00:00:12,680 --> 00:00:12,880
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir à tous les
24
00:00:12,920 --> 00:00:13,080
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir à tous les deux.
Here is the result after using "merge lines with same text" in SE, this is much better but there are still a lot of double lines ;
1
00:00:01,920 --> 00:00:03,240
Heureux de vous retrouver pour ce
2e numéro de la 11e saison de "On
2
00:00:03,280 --> 00:00:04,320
2e numéro de la 11e saison de "On
n'est pas couché" avec notre
3
00:00:04,360 --> 00:00:06,360
n'est pas couché" avec notre
nouveau duo : Vanessa Burggraf et
4
00:00:06,400 --> 00:00:13,080
nouveau duo : Vanessa Burggraf et
Yann Moix. Bonsoir à tous les deux.
I got the idea to use a 2nd time "merge lines with same text" but it doesn't detect the double lines anymore (and the display duration is too long, 7 seconds for n° 4, but I believe I can limit duration).
S_E_New
8th October 2016, 15:10
@Nikse555 is there any way I can do this?
http://i.imgur.com/XywyD0k.png
What I mean is that if there is any way to put subtitles with the same start, end and length at the same time, provided that the first line has an alignment (Top/Left, Top/Center, Top/Right, etc)
:thanks:
Nikse555
8th October 2016, 20:40
(edited post) Are you talking about image based export or just saving in .ASS S_E_New ?
S_E_New
9th October 2016, 07:36
Sorry for not expressing myself properly, I'm talking about image based export (VobSub).
Music Fan
9th October 2016, 11:41
@ Nikse555 : I believe you forgot to answer my question ;
http://forum.doom9.org/showthread.php?p=1779846#post1779846
:thanks:
Zetti
30th October 2016, 21:27
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.0
LouieChuckyMerry
4th November 2016, 19:55
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.0
+1 :thanks:
NikosD
12th November 2016, 20:10
Hello.
Is there any other way besides "teaching" OCR process in order to convert the VobSub subtitles of a DVD to SRT file ?
Unfortunately, I don't know if this is a problem of my Bluray player, but when I feed my Bluray with a MKV file (DVD rip) with embedded VobSub file it doesn't recognize it at all.
Also if I extract the original IDX/SUB subtitle from MKV and put it in the same folder with MKV, I have a lot of problems with my Bluray player, although it can read it and display the subtitles.
If I seek forward/backwards it looses sync of external IDX/SUB subs and even If I don't seek, sometimes the IDX/SUB subtitles are just standing on screen for many seconds after audio finish.
So, I can't see any other solution but converting IDX/SUB to SRT.
Any other idea ?
Any easy way to convert IDX/SUB to SRT ?
P.S
All of my converted DVD rips to MKV work fine with subtitles using PC.
The problem is about Bluray displaying subtitles (VobSub).
sneaker_ger
12th November 2016, 20:23
If your player is broken there isn't much to fix.
Things you can check:
- try embed in mkv with zlip compression (mkvmerge default)
- try embed in mkv without zlip compression
- what is the resolution of the idx/sub? Try to downscale to e.g. 720x480, 720x576 or video resolution. (BDSup2Sub can do that, don't know about Subtitle Edit)
- every combination of the above
NikosD
12th November 2016, 20:30
I have tried both two versions suggested - with or without zlib compression and with uncompressed subtitles the situation is a little better but not a miracle.
I haven't changed the original resolution of 720x576, it's already low I think.
My Blurays 3D are both LG and relatively recent models like BX580.
I have no other problem besides embedded VobSub and external idx/sub rendering.
So, are you saying that other Blurays can handle VobSub or idx/sub in a perfect way ?
Thanks for your reply BTW
Music Fan
12th November 2016, 20:49
You can also convert your subtitles to Blu-ray sup with SE and make an AVCHD with tsMuxer, admitting your player handles AVCHD (on usb key, external hard drive, dvd or BD-R) and if your videos have a standard resolution (that's why I never remove black borders, anyway it doesn't need a lot of bitrate).
sneaker_ger
12th November 2016, 21:07
My Blurays 3D are both LG and relatively recent models like BX580.
BX 580 is over 6 years old. Maybe it's worth to check some mkv compatibility options? Add "--clusters-in-meta-seek --disable-track-statistics-tags --engage no_cue_duration --engage no_cue_relative_position" as additional options into the mkvmerge command-line (in "Output" tab in the GUI).
So, are you saying that other Blurays can handle VobSub or idx/sub in a perfect way ?
I don't really see a lot of complains about it but I couldn't name any recommendations. I know people wouldn't be converting their BluRay subtitles to idx/sub if those were problematic on every player.
Or maybe your subtitles are special (fades/animations?) exceeding specs.
NikosD
12th November 2016, 21:21
My subtitles are not special,they are very simple.
I thought that most people and most "pirates" convert the VobSub subtitles to SRT, not idx/sub.
I will give a give a try to your command line suggestions.
NikosD
12th November 2016, 21:52
You are right, the BX 580 is a 2010 release and my other Bluray is BD670, a 2011 release.
Not so new, but great BDs in their time.
sergio2
18th November 2016, 16:28
My Samsung TV does not read PGS subtitles, so I use to import those subtitles from .mkv files and convert them into .srt files with Subtitle Edit.
I then remux .mkv files with .srt subtitles with MKVMerge to display .srt subtitles on my TV.
The thing is that subtitles wich are well located in downscreen black bar at the begining of the video, move up on the junction between black bar and picture after a while.
Any clue to maintain them in downscreen black bar ?
hubblec4
18th November 2016, 19:13
You could try BDSup2Sub. Convert your PGS into a Sub.idx/Sub.sub (DVD subtitle format) without changing the resolution and position.
locotus
11th December 2016, 06:15
After using V3.50 for a couple of wees, found that when converting srt subs to vobsub the program
forgets selected font color, even without closing the program.
In last multiple replace modification import and export functions were eliminated. There were a nice way
to transfer that data base from one computer to another. If posible, would be nice to recover both of them and include some way of sorting to ease checking and corrections.
Thanks in advance.
Music Fan
11th December 2016, 10:36
@ Nikse555 : I believe you forgot to answer my question ;
http://forum.doom9.org/showthread.php?p=1779846#post1779846
:thanks:
That was 3 months ago, still waiting ...:scared:
Zetti
14th December 2016, 16:45
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.1
Havokdan
14th December 2016, 19:30
My antivirus always deletes the software, note: https://i.imgur.com/SgaeSHO.png
Nikse555
15th December 2016, 16:43
@Havokdan: Try to update your anti-virus definitions...
Nikse555
15th December 2016, 16:55
Yep, SE 3.5.1 is out - contains mostly minor bug fixes and thx to var1ap Subtitle Edit now has a nice video player on Linux, libmpv :)
@Music Fan: Sorry, don't have too much time atm - but I do have a lot of feature requests...
@locotus: The import/export functions has been "hidden" in the list view's context menu, so right-click in the rules list view and choose "Import..." or "Export..."
StainlessS
15th December 2016, 17:14
Havokdan, Just uploaded 3.5.1 to VirusTotal, where 1/56 said malware, trojan keylogger as for your post, that one being nProtect (presume that is what you are using, virus definitions were current today).
Suggest that you submit to nProtect as false +ve.
report here:-
https://virustotal.com/en/file/2e10ca137b5371e8e5ff8b064a9ec6fc4821e0db185e2468109ba4855e1c8322/analysis/1481818196/
Nice one Nikse555, thanx :)
amayra
15th December 2016, 18:56
i have Feature requests is not necessary but just for save time i use this tools but i prefer to use it with SE if you add GUI (like Tesseract ) for this tools this well be great :
1 Sushi (https://github.com/tp7/Sushi)
2 Prass (https://github.com/tp7/Prass)
Havokdan
15th December 2016, 19:36
Havokdan, Just uploaded 3.5.1 to VirusTotal, where 1/56 said malware, trojan keylogger as for your post, that one being nProtect (presume that is what you are using, virus definitions were current today).
Suggest that you submit to nProtect as false +ve.
report here:-
https://virustotal.com/en/file/2e10ca137b5371e8e5ff8b064a9ec6fc4821e0db185e2468109ba4855e1c8322/analysis/1481818196/
Nice one Nikse555, thanx :)
Detail, this happens since the penultimate version. But I'll try again. I am using Kaspersky Internet Security
locotus
15th December 2016, 20:31
@locotus: The import/export functions has been "hidden" in the list view's context menu, so right-click in the rules list view and choose "Import..." or "Export..."
Thanks, hard to find.
But hang while importing or re-writting the old template.
By the way, was tryng to find the templete file a couple of
days ago to manually overwrite or copy the template I exported previously from 3.41.3 to 3.51 but couldn't find it nor in program foulder nor in Appdata/Roaming/subtitle edit.
Thanks.
Thunderbolt8
16th December 2016, 00:32
would it be possible to add the option to save certain subtitle formats & styles for batch conversion? whenever I just want to convert something else than .ass I have to re-specify all .ass parameters manually again after that (font size, colors, vertical margin, order values, video resolutio, wrap style...).
Thunderbolt8
18th December 2016, 15:07
would it also be possible to add options in batch conversion to choose the order of how convert options will be applied? usually its better to use "remove text for HI" before "fix common errors", but in case of some web-dl subs which consist of lots of instances with more than two spoken lines per subtitle line its better to apply fix common errors before removing text for HI, because "remove text for HI" works better if subtitles consist only of two spoken lines per subtitle line and not three (which then gets fixed before applying the HI fixes)
locotus
15th January 2017, 20:41
Hi, think I found a bug in Spell Check, at least in Spanish.
When a word that need to be corrected or change is between interrogation marks (¿?)
it seems to work in the spell check window, but really don't apply the changes.
Need to abort the spell check and do the change manually. Then, run again spell Check.
Thanks.
khoahoc0508
16th January 2017, 03:29
Thanks for update.I'm very much interested in try it on OSX.But i can't find where to download.Please give a link to download or tell me how to compile it :thanks:
Nikse555
28th January 2017, 19:08
@Thunderbolt8: thx for the info - I've tried to fix the clearing of ASS/SSA styles after a format change in batch convert in latest beta.
@locotus: thx for the info - the bug regarding spell check change word and "¿?" should be fixed in latest beta.
@khoahoc0508: You could try mono or wine. I did start on a mac version, but it seemed to be hard as I could not find any .net "image" component (or use System.Drawing components) for mac - nor did I get much feedback on it.
https://github.com/SubtitleEdit/subtitleedit-mac/releases
Latest beta, portable version (install anywhere, but under "Program files"): https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.1/SubtitleEditBeta.zip
Thunderbolt8
6th February 2017, 13:37
thank you
varekai
22nd February 2017, 11:00
Hello,
After editing srt subtitles sometimes I get empty lines.
I can't figure out how to delete all lines in one go.
Now I manually mark each empty line and then hit delete.
Would be great if someone could enlighten me on this?
Thanks...
Edit:
Found a workaround:
Options-->Settings-->Check "Remove blank lines when opening a subtitle".
But I would like to do the remove when I finished editing the srt.
The workaround setting requires to first edit the srt,
check the "Remove..." then close Subtitle Edit and open it again and blank lines will be removed.
I will also have to uncheck "Remove..." before opening a new srt project to have the original srt untouched.
Hope I'm making sense... and many thanks to the creator of this excellent software.
Regards
Boulder
22nd February 2017, 12:53
You could use Fix common errors to remove the empty lines after you have finished working on the file.
varekai
22nd February 2017, 13:52
You could use Fix common errors to remove the empty lines after you have finished working on the file.
Aha! Excellent! That's exactly what I was looking for! Thanks!
Case closed! :D
Regards
Nikse555
3rd March 2017, 16:53
Subtitle Edit 3.5.2 is out: https://github.com/SubtitleEdit/subtitleedit/releases
List view can now display cps and/or wpm + many minor improvements and fixes :)
Lucius Snow
7th March 2017, 11:54
Hello all,
I've got a little problem with Final Cut Pro XML+PNG export. When importing into FCP, all PNG are only visible one frame instead of their real lenght.
However, it works fine in DaVinci...
Any ideas?
Thanks.
hello_hello
3rd April 2017, 18:43
Hi.
Just reporting a bug (I assume) with naming Advanced Sub Station Alpha styles (I'm using Subtitle Edit 3.5.2 on XP).
When creating a style and giving it a multi-word name (ie Magnificent New Style), Subtitle Edit drops everything after the first space, so "Magnificent New Style" becomes "Magnificent ".
Style: Magnificent ,Courier New,20,&H00FFFFFF,&H0300FFFF,&H00000000,&H02000000,0,0,0,0,100,100,0,0,1,2,1,2,10,10,10,1
When saving the subtitle file, for the individual subtitles, it also drops the space.
Dialogue: 0,0:22:54.82,0:23:01.12,Magnificent,,0,0,0,,Subtitle text......
The above only happens when first creating/naming a style. If you then edit the style without editing the name, when you re-save the subtitle file, Subtitle Edit also drops the space from the style name, so the subtitles go back to working properly again.
For Sub Station Alpha subtitles, the space after the first word is dropped as soon as the style is created (or the subtitle file is first saved), so even though the style mightn't have the expected name, it doesn't break functionality.
Style: Magnificent,courier new,20,-1,-50331904,-16777216,-50331648,0,0,1,2,1,2,10,10,10,0,1
Dialogue: Marked=0,0:22:54.82,0:23:01.12,Magnificent,NTP,0000,0000,0000,,Subtitle text......
Thank you for the continued work on Subtitle Edit.
Nikse555
3rd April 2017, 20:00
@Lucius Snow: Sorry, don't know (latest beta has improved drop frame support for Final Cut Pro XML+PNG...). Let me know if you find out what's wrong.
@hello_hello: thx for the info :)
I've tried to make a fix (https://github.com/SubtitleEdit/subtitleedit/commit/e48471cad6722e66773037c7a413c781cd7f2ab5) - included in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.2/SubtitleEditBeta.zip
hello_hello
3rd April 2017, 21:16
@hello_hello: thx for the info :)
I've tried to make a fix (https://github.com/SubtitleEdit/subtitleedit/commit/e48471cad6722e66773037c7a413c781cd7f2ab5) - included in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.2/SubtitleEditBeta.zip
That was fast!
I've given the beta version a quick test, and so far so good.
Thanks.
Lucius Snow
4th April 2017, 13:25
@Lucius Snow: Sorry, don't know (latest beta has improved drop frame support for Final Cut Pro XML+PNG...). Let me know if you find out what's wrong.
The drop frame concerns 29,97 fps frequency. I work at 25 fps so I'm not concerned. I've tried anyway the new beta, no changes.
Music Fan
4th April 2017, 19:14
Hi, still waiting for a response after ... 7 months ;
https://forum.doom9.org/showthread.php?p=1779846#post1779846
Barneyk
11th April 2017, 22:18
Subtitle Edit 3.5.2 is out: https://github.com/SubtitleEdit/subtitleedit/releases
List view can now display cps and/or wpm + many minor improvements and fixes :)
Is there a way to disable the videoplayer and stuff?
It is just annoying to me that the Controls, Audio and Video windows pop up everytime I open a subtitle to do a quick edit.
The program also crashes and locks with many videofiles. And even when it works well it feels slower to open, so I would just love to be able to disable it when I don't wanna use it.
Thank you for the great work. :)
Nikse555
16th April 2017, 11:25
@Music Fan: Sorry, I started on it but did not find a good solution...
@Barneyk: Just had a few requests for that, so it's included in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.2/SubtitleEditBeta.zip - see Options -> Settings -> Video player - "Auto-open video file when opening subtitle"
LDD9O
22nd April 2017, 12:35
Thank you for constantly updating this. It currently my main subtitle software.
I'm not sure if this feature is available yet, the option Word List > OCR seem quite close to what I need but it only get activate during OCR fix up tools.
What I'm interest is, "short key" or "short words".
When I subtitle there tend to be quite many repeat words and phrases (e.g Names or Idiom). Typing these out each time is a pain, I usually Copy/Paste it but clipboard can only hold one. Other method is I use a Acronym and use Replace (Ctrl+H) afterward.
So the feature I'm hoping you add is the ability to use short cut/key and it will auto change to what I need.
Feature: Short Cut/Key Words
Example: Typing "abc" will replace it with, "all can bee", typing lol change it to "laugh out loud". These words and replacement key can be changed of course.
Thank you.
rockydon
22nd April 2017, 14:02
Hello I have been using subtitle edit,I wanted to convert this non-english sub into srt
It's Hindi sub in pgs I want to convert into Hindi srt
I updated my software also updated Hindi dictionary (downloaded as I was ask to do so when I selected Hindi language on OCR main Page after loading pgs)
here is PGS - https://www.sendspace.com/file/2s9jkh
here is SRT - https://www.sendspace.com/file/ascd2t
I got this srt,so pls let me know what to b done
hello_hello
25th April 2017, 11:30
When I subtitle there tend to be quite many repeat words and phrases (e.g Names or Idiom). Typing these out each time is a pain, I usually Copy/Paste it but clipboard can only hold one. Other method is I use a Acronym and use Replace (Ctrl+H) afterward.
While not directly related to Subtitle Edit there's a couple of other options you could try. One is a clipboard manager. Actually, I find Windows quite annoying to use without one. Some clipboard managers let you save permanent clipboard entries. The clipboard manager I use is ClipX (http://bluemars.org/clipx/) because it also lets you configure the type of clipboard items it saves. I have it set not to save images in the clipboard.
Another option might be a macro program. Maybe give AutoHotkey (https://www.autohotkey.com/) a try. One of it's most basic jobs would be to save strings of text to automatically type for you after you type a "hotstring" (pretty much any key combination you choose). As a simple example for a script such as the one below, typing "zzs" would cause AutoHotkey to replace "zzs" with "This is the phrase I don't want to repeatedly type" followed by the enter key.
;Some phrase.
:*:zzs::This is the phrase I don't want to repeatedly type{Enter}
AutoIT (https://www.autoitscript.com/site/autoit/) also seems to be quite popular.
Nikse555
12th May 2017, 19:04
@rockydon: Latest version works on your Hindi subtitle (at least for me). You'll probably need to download the Hindi Tesseract dictionary from inside SE again.
Also, Subtitle Edit 3.5.3 is out - https://github.com/SubtitleEdit/subtitleedit/releases
Partial change log:
* Minor OCR improvements
* New Netflix quality checker - thx askolesov/pavel-belenko
* Added optional list view column "Actor" for ASS/SSA - thx william
* New subtitle format
* Import plain text now also supports input as HTML
* Added "Total words" to statistics - thx Barbara
* Added more Tesseract dictionaries - thx lucybook/Elheym
* Added "Video auto-load" to UI settings (so it's now possible NOT to auto-load videos) - thx mtarini
* "names_etc.xml" renamed to "names.xml" + local dictionaries now has a "blacklist" - thx ivandrofly
* In "Change casing" it's now possible to add extra names
* Export image margins are now percentage - thx aaaxx/gorgorias
* "Import plain text" now rembembers options - thx Leon
* "Insert subtitle here..." added to waveform context menu
* Drag'n'drop of subtitles now allows up to 10 MB - thx Silvena
* Fixed crash in "Add to user dictionary" in OCR - thx alfix0
* Fixed original file name bug in "Undo" - thx darnn
* Fixed non default timecode scales in MKV - thx mkver
* Fix missing UTF-8 tag in OCR HTML export - thx dgonyier
* Fixed "FCP + image" export for drop frame rates - thx chris
* Fixed finding Tesseract traineddata files on Linux - thx mooop12
* Fixed loading of "mks" files from cmd line - thx Charles
* Fixed new ASS/SSA style with space in name - thx hello_hello
* Fixed crash in "Binary image compare" OCR with empty image
* Some minor fixes for spell check "change whole word" - thx N/A
Barneyk
13th May 2017, 16:56
Disabling the "auto-open video file" setting doesn't seem to work for me.
Even though I have it disabled it still auto-opens.
And it leads to subtitle edit crashing about 50% of the time I try and use it. And even when it doesn't crash it makes the program open really sluggishly and it is a pain to use all of the sudden.
Is there a way to disable the videoplayer function completly?
I just wanna edit subtitles, I want nothing to do with the videofiles.
Nikse555
13th May 2017, 19:56
@Barneyk: Thx for the info :)
Could you try latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
(Also, you could try "mpv media player" - you can download the single dll file directly from SE video player settings.)
Barneyk
13th May 2017, 21:26
@Barneyk: Thx for the info :)
Could you try latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
(Also, you could try "mpv media player" - you can download the single dll file directly from SE video player settings.)
Seems to work now but it does still crash with some videofiles with mpv media player, not as many though, most of the ones I tried worked better now when I did load the video.
But it still freezes and have to be terminated with some files.
But it doesn't load the video files now so its fine as far as I am concerned. :)
I would prefer it if there was an option to not open the video windows at all, they are in the way and makes the program feel clunky to me.
I undocked them but they still open everytime I open the program whether I load a video file or not.
sneaker_ger
13th May 2017, 21:39
It's not remembering this switch for you? For me it does.
https://abload.de/img/subtitle_edit_video_-p9u94.png
von Suppé
15th May 2017, 07:35
Would it be possible to get back the bottom offset unit in SUP export for pixels again?
I'm having trouble with precise height adjustment and I loved the pixelamount a lot better than the percentage unit.
cheers
hello_hello
16th May 2017, 08:58
Nikse555,
I have a request for a little tweak, if possible.....
In the OCR import window there's a drop down box to select the dictionary to use for OCR correction. Would it be possible to tell Subtitle Edit to remember the last used dictionary rather than always defaulting to English US? As I'm silly, I'm always forgetting to change it until after I've started the OCR process.
Cheers.
Has anyone got any luck with running SE with video on Linux? I struggle with that, the program runs fine with mono but all video options are greyed out. But I have MPV installed (the linux version).
dulanfai
24th June 2017, 07:23
Hey guys, I need help here. I tried to import a vobsub into SE. And then I deleted 2 lines. When I exported it, the text style is a little bit different. Is there anyway I can keep all the settings from the source? I tried many times, the output text is always slightly different. Even if I didn't delete anything from it. A simple import/export will always change the vobsub file size. I guess SE is auto cropping or auto optimizing the graphic in the vobsub file?
Any way I can actually keep everything same as the source vobsub?
Music Fan
24th June 2017, 09:39
How did you export it ? Did you first make OCR ?
dulanfai
24th June 2017, 09:56
Nah, I just drag a vobsub into SE. And then right click and export it to Vobsub. The settings always not the same as the source vobsub. If there is an update I hope this will be fixed. The vobsub file size is also smaller then the source file. (which suggest it is optimized? I don't know) I use SubResync and do the same thing, the output file size is same as the source and the text style is also same. But too bad SubResync can't delete a few lines as effective as SE.
Music Fan
24th June 2017, 11:10
Try with BDSup2Sub.
dulanfai
25th June 2017, 04:31
Actually, BDSup2Sub is not as effective. I need to delete frame 1 by 1. Now SE is the only software that allow me to delete a few lines in one action. That's why I am asking if there is any work around for this. It is ok. I am now using Subresync.
von Suppé
20th July 2017, 19:35
I know I asked this earlier. I'm having trouble with the newest SUP export percentage settings.
Is it much trouble to bring back (or be able to choose) the good old pixel offset and size units into 'amount of pixels' again?
Nikse555
21st July 2017, 06:23
@dulanfai: Re-saving a vobsub file is not supposed to change the image pixels. The container format and RLE-encoding is re-done though.
@Von Suppé: I think it's possible to choose between pixels and percentage in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
von Suppé
22nd July 2017, 08:01
@Von Suppé: I think it's possible to choose between pixels and percentage in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
This would be great! I go try.
Thanks in advance, Nikse555 :thanks:
Edit: First quick test shows it works like a charm. Me a happy fellow :-)
So much thanks for your trouble, Nikse
Working with Advanced Sub Station Alpha (*.ass file), in the ASSA style window, the opacity (Alpha channel) of "Outline" settings is honoured in the preview. The opacity of "Shadow" settings is not honoured in preview.
Is this a bug?
von Suppé
23rd July 2017, 08:36
Hi Nikse,
I'm still encountering problems with .ass files exporting to SUP. SE still doesn't remember line-height settings.
In the SUP export window the Line height is default 73. I select all lines and change to 62. Then when I select one line, the height is back to its default 73. Things get more confused when I only want to select one line or lines with the same style.
Same goes for Alpha channel setting and Shadow width. Different styles sometimes ask for different Shadow, Alpha and Line height settings
Also, when possible, in the SUP export window, I'd like to have a feature that allows me to select all the lines with a specific ASSA style. I sometimes use quite a lot of different styles within the .ass file and then manually selecting the lines with the same style can be a tedious job.
Cheers
von Suppé
Nikse555
25th July 2017, 10:27
Yes, the shadow opacity looks like a bug - thx. I've tried to fix it here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
The line height works for selected "style" (style name from first selected line is used and style name should be visible below the height input box)
Rack
29th July 2017, 07:27
Hi,
Is there a way to extract DVB Subtitles using Subtitle Edit directly without first using ProjectX or CCExtractor to extract the images and then OCR them with Subtitle Edit?
Music Fan
29th July 2017, 09:04
Yes, open the TS file, it will detect the DVB Sub.
You can also export it directly in Blu-ray sup, without making OCR (right click on text in OCR box, export). But the size is much more little after exporting in Blu-ray sup. Nikse555 could not explain this when I exposed the problem 2 or 3 years ago.
Rack
29th July 2017, 19:00
Yes, open the TS file, it will detect the DVB Sub.
You can also export it directly in Blu-ray sup, without making OCR (right click on text in OCR box, export). But the size is much more little after exporting in Blu-ray sup. Nikse555 could not explain this when I exposed the problem 2 or 3 years ago.
Actually, I've tried that before, but SE didn't recognize the file.
So, I re-saved the file as .ts again using VideoRedo and tried with SE once more and SE recognized it this time.
The issue now is that there are 580 subtitle images while I'm getting 1160 images in SE.
After each subtitle image, there is an empty image.
To be fair, I faced the same issue with ProjectX and CCExtractor and that's why I wanted to try SE.
Nevertheless, when I compared both subtitles, the embedded DVB subtitle in the .ts file and the subtitle resulted from SE using PotPlayer, both of them where almost in sync, except for a delay of 1-2 seconds throughout the whole video.
Nikse555
30th July 2017, 16:11
@Rack: Could you possibly email me the first 1-20 mb of the original file that SE does not recognize?
Rack
30th July 2017, 18:14
@Rack: Could you possibly email me the first 1-20 mb of the original file that SE does not recognize?
Hi,
Regret the inconvenience. It seems that I only needed to change to extension of the file from .nds to .ts and SE accepted the file without any issues. Media Info is already showing that the file type is: MPEG-TS.
von Suppé
21st August 2017, 11:17
Yes, the shadow opacity looks like a bug - thx. I've tried to fix it here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
The line height works for selected "style" (style name from first selected line is used and style name should be visible below the height input box)
Thanks Nikse, alpha box is greyed out now, but I'm encountering something new in the beta.
When I select first line and then go down to the second (with mouse or arrow down) SE beta freezes?
Nikse555
22nd August 2017, 18:12
Thanks Nikse, alpha box is greyed out now, but I'm encountering something new in the beta.
When I select first line and then go down to the second (with mouse or arrow down) SE beta freezes?
Hm, I cannot re-create this... could you email me the sub? (I presume it's only with one special subtitle?)
von Suppé
23rd August 2017, 07:50
...I presume it's only with one special subtitle?
I'm sorry, no it's not. I presumed we both realized that it's about working with more than one ASSA style. Sorry if I wasn't clear enough earlier.
I have created a small .ass file for testing purposes with (for now) two different styles/fonts and positioning.
One more thing, can I unwantedly import problems when the test-file (.ass) is created in Aegisub? My testfile comes from it because of the realtime preview in this software.
Thanks in advance
von Suppé
24th August 2017, 13:36
Has anyone ever tried to edit a XML+PNG-file in SE?
I mean: OCR it --> edit in text-editor --> then export to XML+PNG or SUP with honouring the X and Y offset-info that were in the original xml-file? (at least for the lines that are unchanged, and remove all data for lines that were deleted during edit?
xKupo
26th August 2017, 20:32
Hi,
I have a problem with the spell check and auto fix. For some reason SE changes the first letter of certain german verbs from like a to A, b to B and so on. And I just do not Understand (yea like this) why it does what it does.
I noticed that it's because of the 'Auto fix names where only casing differ'-function but my word lists are clean. The changed verbs do not appear in that lists.
It's annoying because it rarely happens so I just noticed it after like 80 subtitles and I do not have the idx files anymore.
The main problem is that I don't know what words have been changed. If I find out for what reason SE changed these words I maybe find the other changed words too.
So does anyone have an idea for what I need to search? Any certain rules the auto fix is following?
Boulder
27th August 2017, 16:28
What was the reason for removing OCR via image compare? I used it all the time without any issues on both VobSub and Blu-ray subtitles.
Nikse555
31st August 2017, 07:50
@xKupo: SE probably changed the "names" because your spell check dictionary only contains an uppercase version of the word.
@Boulder: I don't see any reason to keep "Image compare" (a lot of old code) and ppl might start using that instead of the newer "Binary image compare".
I think the "Binary image compare" OCR methods is better than "Image compare" - could you verify with your files?
Boulder
6th September 2017, 18:59
I tried it, and after a few rounds it seems to be working better. There were some issues with it detecting a hyphen as a colon for some reason and some other random things, but re-teaching them and setting the max error percentage very low (2 %) apparently fixed them.
Would it be possible to make the program remember the last selection? Now it's always offering Tesseract since it seems to be the first in the list.
Nikse555
6th September 2017, 19:43
@Boulder: thx for testing :)
Let me know if some of the defaults should be changed - and I might improve some detection if you e-mail/link the sub/image.
The remember OCR method should work in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip (should be out in 3.5.4 final soon)
von Suppé
7th September 2017, 09:35
...(should be out in 3.5.4 final soon)
Any chance that the issue I spoke aboute here: https://forum.doom9.org/showpost.php?p=1815953&postcount=542
will be addressed in latest final, Nikse?
Nikse555
8th September 2017, 15:07
@von Suppé: If I can re-create this, then I can probably fix it... how do I re-create this in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip ?
von Suppé
9th September 2017, 07:18
I'll go testing. Will let you know.
Thanks in advance, Nikse.
I have here xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx
a small .ass file with two fonts. Loading it up in latest beta. Directly go to export --> bluray SUP, SE freezes
Edit: take this file https://nofile.io/f/qpGWVHh55qx/Segoe+Script+and+AgencyFB.ass
previous file had corrupt outline-pixel numbers (no whole numbers; apparently Aegisub can put that out??))).
Line height is not remembered and not shown in preview. Next to that font height is much much larger than preview in Aegisub, hence my previous question: Can I expect unwanted errors when imporitng .ass file that is created in Aegisub?
mbcd
9th September 2017, 19:20
*lol* Cant hold this for myself ... :p;):D
You have lots of "X", a small ".ass" and two "fonts" ???
Youre a woman ?
What size are those two "fonts" ? D-Size or DD ?
von Suppé
10th September 2017, 09:41
:D:D:D Ooooooh yeah, rub it in!!! :D:D:D And you haven't seen the footage yet the subs belong to... :eek::eek:
Maybe Nikse will be so kind to create the possibility to "save as..." custom-made "DD-Cinema" format for you?? LOL
Nikse555
11th September 2017, 06:54
He, I do wonder why ASS was chosen as file extension - maybe to keep the format a bit underground ? ;)
@von Suppé: thx for the file - new beta upped: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
jpsdr
11th September 2017, 08:57
Original name is SSA : Sub Station Alpha. When they improved it, they just chose for the improved version the reverse name ASS, for Advanced Sub Station.
Well in case your question was a real question... ;)
von Suppé
11th September 2017, 10:46
A quick test with your latest beta still show problems, Nikse.
After opening and directly go to --> export as BD SUP, in the list window all (7 in this case) lines are shown in "Segoe Script Red shadow alpha 160" font. This makes it hard to choose another style to select and edit the lineheight from.
The size in the preview looks more reasonable now, but... read further.
In this example test, changing the first line's lineheight goes very troublesome, direct keyboard input reacts erratically crazy? It does however remembers the lineheight for different styles, it seems. However, the preview doesn't react on changing lineheight though and that is a pitty; "I don't see what I get."
After that, I can't change the lineheight of the second style with direct keyboard input whatever I try and I'm stuck with the up/down arrows.
After clicking "Export all lines", it seems SE hangs and output SUPfile goes incredibly slow. First I thought it was freezed. No change in progress bar. After output finished, the preview window shows a change-back to abnormal big size again. (Video resolution settings was 1920x1080).
SUP ouputfile, opened with BDSub2Sup shows too big subs. First line is 1767x317 pixels. Line 6 is 1920(!)x354 px and is cropped.
Nikse555
11th September 2017, 13:15
OK, new beta up: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
The list view font issue should now be fixed - also shows style name now as that might help.
The export "hang" should also be fixed (or at least better).
Font size... OK, in latest beta SE now just uses the raw size - but I don't know if that is better.
The no-preview-stuff I cannot re-create.
von Suppé
11th September 2017, 20:50
Input issues are solved. Very nice to see the proper style AND "Style column" in export listview! Big improval, thanks.
Output SUP is faster, but still shows too large subtitles, same as first test. I wonder if the font size unit in Aegisub is in pixels.
...in latest beta SE now just uses the raw size...How must I interpret "raw"?
Still, no change in preview when changing the line height though. Maybe this is baked in the style in Aegisub and the same reason that the fontsize still has issues?
The differences between Aegisub's preview and the SUP output of SE are HUGE.
Edit: My believe that something is wrong with the way your latest beta is handling both the preview in the SUP export window and the SUP output itself, grows stronger. In SE 3.5.3 the preview flawlessly reacts on lineheight changes. SUP output still gives too big lines, but my guess is that this still is a size-interpretation issue between Aegisub and SE.
When downsizing the ASS Style fonts in SE 3.5.3, the quality in the preview (SUP export window) stays good
In latest beta however, changing fonts to same size, the quality is bad, like SE hasn't pixels enough to create proper quality.
Nikse555
12th September 2017, 09:17
New beta up again: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
I've tried to calculate an equation between ASSA (no ass here) font size and the .NET text render font size and the result is hopefully that the font sizes matches better?
(If this works I'll have to do the same for "Simple rendering")
Do you still have problems with the preview/input? Perhaps a screenshot can show what the problem is?
Nikse555
12th September 2017, 17:40
Also, if you use libmpv as video player in SE then subtitle preview (even ASSA) on the video should be displayed correctly (by libmpv).
Ghitulescu
12th September 2017, 19:43
I read the latest log, am I right to assume some improvements in supporting true BD subs have been achieved?
For ASS/SSA/SRT there are zillions of software supporting them...
Nikse555
13th September 2017, 07:32
I read the latest log, am I right to assume some improvements in supporting true BD subs have been achieved?
Probably not... what do you mean by "true BD"?
SE can import/export BD sup + BD TextST (import also from ts/m2ts/mkv/mp4).
von Suppé
13th September 2017, 10:09
Nikse, I found out something.
Before you furtherly go put a lot of effort in SUP output size, or waste tons of batteries :D calculating some size-equation between Aegisub and SE; this was never an issue in SE 353!
After some comparison tests, and having read back a few pages I realized that the SE 353 problems I initially posted were about the remembering of lineheights and some more. Because I'm so busy focussing on these issues, I plainly overlooked the fact that the SUP-preview, -quality and -size problems were introduced in later beta's. I was stunned, really... and I deeply apologize for the trouble you have taken to solve this non-problem :(
I could have prevented it if I had payed more attention when the non-issue became an issue when beta's came out.
So, keep whatever formula you had in SE 353 concerning the SUP quality & size.
But, when releasing a new version, please do add the better features the latest beta does have concerning:
1 - proper remembering of Line heigths for each ASSA Style
2 - proper list-view of different styles AND the "Style column" itself (these are great!)
3 - the Alpha setting, border- and shadow colors are greyed out in latest beta. As it should, my guess, as these features are defined bij ASSA style itself.
4 - posibility to chose pixels or percentage in bottom- and left/right margins (when working with srt, that is)
One thing concerning the SUP output still:
After a quick test with one ASSA Style only (because of the heigth setting remembering issue) and the same subtitle as SRT file, I experienced that the sizes of the subtitle windows (checked with BDSub2Sup) were the same, but some differences in positioning take place. I think this has to do with the "Left/right margin" settings, that are not greyed out in the SUP export window. As the horizontal and vertical margins are defined within an ASSA Style, maybe it must be "disabled"? I am not sure if these offset values do the same in Aegisub, compared to SE though...
Nikse555
13th September 2017, 13:06
@von Suppé: I think I'll keep the current beta sizes... fits somewhat with what mpv/vlc shows.
About the quality stuff - I cannot see anything here and I cannot see any reason why quality should be changed.
thx about the "margin" settings - was not taken from ASSA, but this should also be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
Ghitulescu
13th September 2017, 15:13
Probably not... what do you mean by "true BD"?
SE can import/export BD sup + BD TextST (import also from ts/m2ts/mkv/mp4).
These are "true BD" subs, not various formats mostly based on SRT or SSA ...
That it can import (then even process them eg via OCR) is known to me since a very old version I use from time to time.
But my old version rearrange them (it disregards the position on screen) and also does not change the colours but rudimentary ...
Anyway, thank you for this software...
von Suppé
13th September 2017, 16:35
Thanks for the latest beta, I'll go test again.
fits somewhat with what mpv/vlc shows.Do you mean during playback in the video-window? I have always experienced the same fontstyle, color and size, whatever Style I use.
Edit:
I think I'll keep the current beta sizes...I would highly regret that. When creating a style in Aegisub, the preview now approx. matches the way it looks in SE 353 preview. New beta preview doesn't look like it. Changing it like latest beta's has an overall effect on how the font looks. I think this has to do with ratio between the size itself and the numbers used in the "Outline" and "Shadow" boxes. I would hate to create an ASSA Style in Aegisub to look good, and then have to edit those styles in SE to meet my initially intended looks.
By the way, I found out that when opening an *.ass file and saving it again as [othername].ass file, I am unable to open that [othername].ass file with Aegisub. Whereas the other way around is going ok. Aren't .ass files supposed to be interchangeable in SE <--> Aegisub both in both directions?
Nikse555
13th September 2017, 19:47
And... a new beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
Border/shadow code back to 3.3.5 build 1 release - better?
Thx for the info about the file not loading in Aegisub... was simply a bug in SE (could not handle [Events] being last). Should also be fixed in latest beta.
von Suppé
15th September 2017, 08:37
Thanks again, Nikse. I'll go testing again.
Cheers
Edit: First tests look good, thanks. During more verbose tests I found out that with SUP export the right margin value in ASSA style is off.
I created some styles with left - & right margins at 15. Left is okay, but right margin comes out too large? (Video reolution set to 1920x1080)
Music Fan
18th September 2017, 15:20
@ Nikse555 : when opening a TS containing 2 DVB-SUP, the parsing is done and we have to select one sup, then the OCR window appears. But after exporting it in Blu-ray Sup, we can't select the other sup without re-parsing the whole TS.
Could you add the possibility to select the sup's from the OCR window (or let the selection window opened) to avoid a new parsing ?
Thanks !
Nikse555
18th September 2017, 18:33
@Music Fan: Hm, I could add a "Save as..." button like in the DVD rip menu. Test version here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip (it saves as raw PES packets - don't know if other programs will be able to handle this).
@von Suppé + any ASSA gurus: I've added support for right margin from ASSA files (used left margin earlier, beta above includes this fix), but now I'm even more confused about the ASSA font size.
If PlayResX/PlayResY is defined - how does that influence the font size? In your "Segoe Script and AgencyFB.ass" file the res is set to 1090x1080 which clearly makes the font size smaller. Any help is much appreciated ! :)
jpsdr
19th September 2017, 09:49
In SSA/ASS all positions are absolute according the video size defined within the file with PlayResX & PlayResY, but became after relative according the size of the video frame. Of course, if size of video frame match the size defined within the SSA, the positions become also absolute.
Rules are following :
If no resolution is set, video is considered PlayResX=384, PlayRexY=288, finaly SRT is considered as the same thing than SSA with no resolution set.
If only PlayResX is set :
- If value is 1280, PlayResY is set to 1024.
- for any other value, PlayResY is set to 3/4*PlayResX.
If only PlayResY is set :
- If value is 1024, PlayResX is set to 1280.
- for any other value, PlayResX is set to 4/3*PlayResY.
So, any "size" thing related is relative to PlayResX/Y. If your PlayResX/Y is set to 640x480, and a font size is set to 30. If your video is 1280x960, your font size will be adjusted to 60, on the other hand, if your video is 320x240, your font size will be adjusted to 15.
I don't know the rule for a change aspect ratio. If your video is 1280x720, i don't know if you adjust your font size by (1280/640) or (720/480) or the average or anything else...
There is no complete document spec for ASS/SSA (there is things, but not complete), the answer i've always found was finaly "the full spec for ASS/SSA is the code source of vsfilter"... :(
Music Fan
19th September 2017, 10:02
This reminds me the problem I have with DVB-SUP whose size is not the same after extraction by SE. When extracted to Blu-ray Sup and remuxed with TSMuxer (which doesn't convert sup), the subtitle is much smaller than when the original ts is played with VLC or MPC-HC. But I don't recode the video, thus the sup should have the same size.
@ Nikse555 : thanks, I will test it as soon as I record another video including 2 DVB-sup.
jpsdr
19th September 2017, 10:05
sup/sub/idx is even more troublesome that it's a layer picture displayed over the video, it's not text based, so, if you don't resize the picture, it will be smaller.
Music Fan
19th September 2017, 10:10
Why ? A simple extraction and format conversion without resizing should not induce size change.
jpsdr
19th September 2017, 11:34
My bad, i've misunderstood what you wrote. One test you can make in that case is to extract the sup also with eac3to for exemple, and compare the TSMuxer result with the sup extracted with SE. But indeed, a sup extracted and remuxed on a video with the same size it's extracted from should produce the same result.
Or maybe... i don't know exactly what is "DVB-SUP", so maybe my assumption is wrong... Is it a way to call a .sup extracted from a Blu-Ray ?
Music Fan
19th September 2017, 11:47
No, it's a sup extracted from a ts recorded with a STB (VU+).
von Suppé
19th September 2017, 19:51
...In your "Segoe Script and AgencyFB.ass" file the res is set to 1090x1080...I assume this is a typo??
Script properties in Aegisub states 1920x1080
Thanks for latest beta, will go testing again.
Music Fan
20th September 2017, 10:59
edit ;
@ Nikse555 : I tested the pes saving option, there is still a parsing when opening the ts, how to use it ?
If I open the pes file, it opens directly one of the subtitles and the timecodes are not good.:confused:
Instead of a "save as" button, could you make an automatic saving for ts parsings to avoid to save it each time ? Or add an option in the settings to activate it or not ?
Crash82
26th September 2017, 01:16
I'm hoping it's okay I'm asking
here I have a small issue in SE
I'm getting this now and then "<i>italic</i>... <i>?</i>" on one line in ocr and I don't know how to make it look like <i>italic...?</i> beside going through each line every time and fix it.
And it does not seem that fix common errors it's able to fix it either and I don't know if it's possible to manually add something to that function somehow.
I also looked into Multiple replace function but I'm not using that functions that much so I don't know exactly how to use it, but I see it's possible to add a regular expression so maybe that will work but I don't know anything about that, I have tried to Google it but I'm not sure what I'm looking for
Hope someone can help me
Thanks in advance
-----------------------------------update-----------------------------------------------------
After I posted about my issue I continue doing some testing
and of course I found out a way to fix my problem.
I don't know if any of you guys have try to have problems for days you can't fix
as soon as you're asking for help you'll find a solution.
I just added to the Multiple replace </i>... <i> to ...
von Suppé
27th September 2017, 10:07
Never mind, sorry...
Nikse555
27th September 2017, 20:18
I assume this is a typo??
Script properties in Aegisub states 1920x1080
Thanks for latest beta, will go testing again.
Yes, that was a typo?
Never mind, sorry...
?
@Music Fan: OK, good catch with the time codes... I guess that will not work then. What do you do with the TS image subs - resave or ocr?
Music Fan
27th September 2017, 22:19
@Music Fan: OK, good catch with the time codes... I guess that will not work then. What do you do with the TS image subs - resave or ocr?
Only resave in Blu-ray sup.
Thus you believe there is no way to save the parsing or let the parsing window opened after selecting one of the sups to select the second after dealing with the first ?
von Suppé
28th September 2017, 11:10
Hi Nikse
Been busy, sorry for late reply.
Latest beta's right offset value still doesn't work proper. Output is (still) too much to the left.
Left offset has stayed okay, though.
The 'Line height' value reacts weird IMHO.
You can imagine, when filling in this value, one tends to imagine a virtual frame around the line(s), in which the subtitle would fit. It would sort of represent the outsides of the rectangled image created (when exporting to SUP or XML-PNG).
Since you have altered the SUP size calculation, the Lineheight value doesn't behave in ratio to the used fontsize number, it feels.
To get my wanted space between lines, I now must use a number that often would be smaller than the fontsize number itself. This feels awkward and not logical. Did you take into account the function of the lineheight value, when you changed the image output size?
Furtherly, I assume that this Lineheight value (besides the space between two lines) also has influence on the amount of "extra, transparent" pixels around all 4 sides of the subtitle. So, IMO, it would somehow co-determine the total rectangle size of the subtitle-image. I hope you understand what I mean with this, it's a bit tricky to explain...
Now, in the Blu-ray SUP export window, would it be possible to implement a frame around the subtitle that would real-time and actively reflect the proper total image with the current Lineheight? I don't mean the preview after "Preview" button, but directly in the SUP-export window itself.
A reason for me wanting this, is because of the result when using (all four, but mainly the bottom -) offset settings.
You can imagine, seeing directly the total subtitle-image, one can make out for himself if adding or reducing offset pixels is necessary.
Thanks for your hard work,
cheers
Nikse555
28th September 2017, 20:29
@Music Fan: Ah, perhaps a context-menu will solve it? Try to right click on a track in the choose subtitle window in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
@von Suppé: thx for testing again :)
The line height calculation seems to work fine for me - can you supply an example where it works poorly?
Furtherly, I assume that this Lineheight value (besides the space between two lines) also has influence on the amount of "extra, transparent" pixels around all 4 sides of the subtitle. So, IMO, it would somehow co-determine the total rectangle size of the subtitle-image. I hope you understand what I mean with this, it's a bit tricky to explain...
cheers
Yes, SE calculates a height to use as a base-line in order to avoid too much jumping up and down of letters.
I guess it would be nice to be able to see this (and if it's too large/small)
Music Fan
28th September 2017, 22:44
@Music Fan: Ah, perhaps a context-menu will solve it? Try to right click on a track in the choose subtitle window in latest beta
There is no change compared to previous version :o : once I selected a subtitle, this window closes, the OCR window opens and when I have finished with this subtitle track, I have to make a new parsing to select to other subtitle track.
Nikse555
29th September 2017, 05:08
@Music Fan: Did you re-download the beta?
http://www.nikse.dk/se-ts.png
Music Fan
29th September 2017, 09:27
Great, you rule man ! :thanks:
Actually the beta version I downloaded yesterday could already do this but I hadn't understood what you meant with the right click.
If I need to OCR both tracks, I can simply begin to export them in Blu-ray sup to avoid 2 parsings, then I can load each sup separately.:)
There is still the size problem, we already discussed about it : after conversion in Blu-ray Sup (keeping 1920x1080 resolution) and remux in ts with TSMuxer, the subs are more little than in the original ts with the DVB Sup, look at these caps (click to get full size) ;
original :
http://nsa39.casimages.com/img/2017/09/29/mini_17092910483243629.jpg (http://www.casimages.com/i/17092910483243629.jpg.html)
after conversion :
http://nsa39.casimages.com/img/2017/09/29/mini_170929104943550113.jpg (http://www.casimages.com/i/170929104943550113.jpg.html)
If you look carefully, you will notice that the original subs are not as sharpened as subtitles that we find on Blu-ray's, as if their resolution was not in 1920x1080 but contained an additional information to be displayed in a greater resolution.
And SE probably detects only the actual resolution without the additional information (a sort of flag).
Or there are well in 1920x1080 and SE detects a lower resolution for some reason :confused:
Thus, after this conversion, I always resize sups with BDSup2Sub (but I could maybe do this with SE, never tried).
von Suppé
29th September 2017, 11:06
Hi Nikse,
Take a look at the 2 screenshots from SE 353 and SE beta export windows. Settings are so, that subtitles have same space between the lines and look the same. Same Lineheight values are used, whereas the fontsizes are differently chosen, because of the resizing you adapted.
SE 353 settings
https://preview.ibb.co/cOjYBG/export_SE353.png (https://ibb.co/fCVjJw)
SE beta settings
https://preview.ibb.co/dR1ryw/export_SE_beta.png (https://ibb.co/mtYh5b)
To be sure, I compared the two resulting .png files.
My findings:
Within the png's, vertical sizes of the text-lines themselves are the same (lucky shot on the different fontsizes :D).
Also they have the same spaces between the lines. Now, IMO this is (next to the fluke that vertical text size is the same) because of the same Lineheights being used.
The widths of the texts are approx. the same.
H/V number of pixels of the total .png files vary a little. I found out, that this is also because SE353 puts out more transparent pixels around the text than latest beta.
Can you imagine, that the ratio of lineheight/fontsize in SE353 is more logical to a person? Lineheight (46) is a bigger number than fontsize (37); this makes sense.
For the same output, in SE beta I had to choose a fontsize value (52) which is higher than the Lineheight (46). This feels awkward in the workflow.
can you supply an example where it works poorly?It is absolutely not working poorly. It does exactly what it needs to do.
I think it is about the auto-interpretation that the units of Lineheight and Fontsize should be the same and that the result (also in preview) should logically and in ratio follow the changes made.
As I earlier said, you can imagine one likes/tends to imagine a "Lineheight-settings dependent" frame around a subtitle, in which it would fit. Lineheight being smaller than fontsize makes no sense then and don't feel good.
von Suppé
29th September 2017, 12:12
@ Music Fan:
Sorry, can't help it, but what does happen to the spoon in the following frames? :D:D
You need subs.....?
Music Fan
29th September 2017, 13:53
The spoon is getting hotter, for sure :D
Nikse555
1st October 2017, 14:37
@von Suppé: I've done some more tests with font size and the beta has been updated with a more precise font size match with vlc/mpv - now the font size is multiplied by 0.895.
Beta link: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
I don't think there's a relationship between font size and pixel size...
@Music Fan: I still believe that the player just upscales the ts-subtitles.
Sorry, no resize for images in SE atm...
EDIT: The above beta will probably be (or very close to) SE 3.5.4 - so please test :)
Music Fan
1st October 2017, 19:49
@Music Fan: I still believe that the player just upscales the ts-subtitles.
I'm not sure about this because on my STB (VU+), the subtitles with the original TS do have the same size than on my pc with MPC-HC.
Thus if MPC-HC makes resize for these DVB-SUB, my STB also does. At the very same size, strange.
Why would this format have to be encoded with a lower resolution than the video and be resized ? This does not make sense to me.
And if it's the case, the information about the resize has to be written somewhere in the stream.
As MPC-HC and VLC do the same resize, you could maybe find informations about this in VLC's codes.
Nikse555
1st October 2017, 20:30
Yes, I guess it's possible some scaling info exists somewhere in the ts file - if anybody knows about this, please do share!
I guess the VLC source contains this info, but it's probably not super easy to find - also I'm not very good at reading C code.
SE now includes the option to resize in the export window: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.3/SubtitleEditBeta.zip
Music Fan
1st October 2017, 21:24
Great, you didn't lose time !
240 % gives good results.
But strange thing : with the resized sups, some lines are grey, others are white, while they all look similar in the non-resized sup. :confused:
Other suggestion : could you add synchro support for sups ?
If no there is a trick which is not very practical : in the OCR window, push on OK without making OCR, the timecodes appears without text. Make synchro, save, re-load the sup and in the OCR window, right click, import new timecodes, right click, export, Blu-ray sup, export all lines.
von Suppé
2nd October 2017, 08:24
The above beta will probably be (or very close to) SE 3.5.4 - so please test :)
Will do! Thanks Nikse :)
Edit: Nikse, nowadays I use export to XML_PNG a lot. It appears that this export mode has a timecode bug?
Seeing back the result in BDSup2Sub the timings are off (at least for 23.976 fps).
Whereas the SUP export timings always were okay (within 1/1000 sec), now in latest beta the timecodes are more off (approx. 45/1000 sec)??
SE353 suffered the same XML_PNG bug btw.
Zetti
2nd October 2017, 23:13
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.4
von Suppé
3rd October 2017, 19:52
Hi Nikse,
Thanks for your hard work :thanks: I'll go test 3.5.4 extensively.
Quick test still shows timings bug with XML_PNG export, at least for 23.976 framerate. I guess you didn't read my yesterday's post yet before launching latest. It sure looks famaliar to BDSup2Sub's bug.
varekai
4th October 2017, 15:32
Having a weird issue with displaying video and audio.
When selecting
Options -->
Settings -->
Videoplayer -->
Video engine -->
Selecting DirectShow I get video but no audio.
Selecting MPC-HC I get audio but no video.
I can't figure this one out.
Probably got something to do with codecs.
I got LAV Filters, MPC-HC, PotPlayer installed.
No extra (evil) codec packs that I know of.
Any ideas how to solve will be much appreciated.
Edit:
Last thing I wanted was to install another mediaplayer but that's what I did.
Installed VLC 2.2.6 win64 and now I got both video and audio...
It's a workaround for now, strange thing is it used to work with MPC-HC alone.
mbcd
4th October 2017, 22:50
Beside of my study, I am working on a new app to handle this feature.
I cant promise you anything, it will take still a long time until I get a release of anything that is worth its name "program" ... because only very less time free to work on it.
Problem is, I cant test anything, because I dont have such a file that contains those multiple text-items.
Is here anyone who could provide me a sample sup-file, please ?
Then I could at least test if my parser is working fine with it right now.
And again: It will take a long time until I get something managed to release.
For first, I plan to do basic stuff, like convert to bdn-xml (both directions), extracting forced, merge normal + forced, and as its best I will implement a way to use a cutlist, so you can adjust a video you edited in an Video-Editing-Software, using the EDL-file-format as cutlist.
So mostly basic stuff that old BDSup2Sub did, but with bugs. It will be forced to work only with bluray-subtitles, not the oder DVD (at least for exporting, not sure about importing).
This are my plans for commandline-options.
Nikse555
5th October 2017, 22:08
@Music Fan: Don't know about the color in the resized subs - seems okay to me, but I guess SE might make the inner color darker, if the sub has a lot of dark outline/shadow. If you have an image sub that's very bad I could take a look.
About the sync in export - would just "move" be okay - like this https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip - or what did you expect/need?
@von Suppé: Could you explain more about the 23.976 problem? Perhaps a source file, and a target file with the expected result?
Do you mean to multiply all times with 0.999 ? Drop frame is annoying... ;)
@mbcd: You can probably find the examples you seek in the BDSup2Sub thread. Is your program written in C# ? On github?
@varekai: I'm not sure about what's the problem is with codecs... you could try to re-install latest version of LAV filters. Also, libmpv is an excellent media player and can be installed from inside SE via Options -> Settings -> Video player.
Normally I prefer DirectShow or libmpv as the seeking is very precise unlike VLC.
Music Fan
6th October 2017, 08:50
@Music Fan: Don't know about the color in the resized subs - seems okay to me, but I guess SE might make the inner color darker, if the sub has a lot of dark outline/shadow. If you have an image sub that's very bad I could take a look.
Here is the sup I use for tests (originally in DVB-SUB format -from a TS- converted to Blu-ray sup with SE, no resize) :
https://www.sendspace.com/file/7bi46s
The same after resize (240 %), the first line looks darker than the others ;
https://www.sendspace.com/file/a238s4
About the sync in export - would just "move" be okay - like this https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip - or what did you expect/need?
I don't see move or anything related to synch for sup in this version :confused: Where is it ?
von Suppé
6th October 2017, 09:12
Hi Nikse
See the attachment. It's a 4 line srt file. For convenience purposes, each line has it's start time, end time, and duration as text.
For people to be able to follow I'll put SE's list view:
https://preview.ibb.co/bKWBMG/SE354_list_view.png (https://ibb.co/iXc7vb)
Then export --> xml/png. Set framerate to 23.976 --> Export all lines. When opening the xml/png result in BDSup2Sub, timings get more off further in time. I took 4 screenshots from BDSup2Sub and put them in one image:
https://preview.ibb.co/jkj98w/4_x_BDSup2_Sub.png (https://ibb.co/dTOfFb)
I don't know if you use BDSup2Sub internally, but it has a framerate bug. r0lZ, the creator of BD3D2MK3D, informed me about this in his software's thread. You can read more here (https://forum.doom9.org/showthread.php?p=1812419#post1812419).
A way to force BDSup2Sub (the standalone I use) to output correct timing info is, in "Conversion Options" window, to tick "Change frame rate" and set both "FPS Source" and "FPS Target" to 23.976. See screenshot:
https://image.ibb.co/muNwow/BDSup2_Sub_Conversion_settings.png (https://imgbb.com/)
It has something to do with a default framerate that the program would assume, so my guess is that it's not a "specific 23.976 fps"-bug, but more framerates suffer from this.
I will test further for SUP output timings.
For now, cheers :thanks:
varekai
6th October 2017, 09:58
@varekai: I'm not sure about what's the problem is with codecs... you could try to re-install latest version of LAV filters. Also, libmpv is an excellent media player and can be installed from inside SE via Options -> Settings -> Video player.
Normally I prefer DirectShow or libmpv as the seeking is very precise unlike VLC.
Yay!! That was spot on! Uninstalled/reinstalled LAVFilters and violá, got DirectShow working with audio and video!
Also cleaned out all I could find that's related to using other mediaplayers and software.
I have som helper apps that BD-RB uses so I couldn't install the latest build of LAVFilters.
There could be some audio sync issues with BD-RB if using any release over LAVFilters 0.65.
Thanks for the info about VLC, havn't been using it for many years, so that is also uninstalled now.
Thank you so much for your help and thanks for a fantastic software!
Best regards
von Suppé
6th October 2017, 15:55
Hi Nikse,
When working with ASSA, would it be possible to implement adding "Actor" when you go Tools --> Sort by?
mbcd
7th October 2017, 09:31
@mbcd: You can probably find the examples you seek in the BDSup2Sub thread. Is your program written in C# ? On github?
Thank you very much for that hint, I found something to test with.
Yeah, it will be in C#, but not on github, dont know what I want to do with it later.
tormento
8th October 2017, 07:50
Am I the only one to get windows form exceptions with 3.5.4 when fixing common errors on srt?
Everything is ok on 3.5.3.
Boulder
8th October 2017, 09:50
I've been using fresh git builds for some time and have had no issues like that.
Nikse555
10th October 2017, 19:28
@Music Fan: I've tried to fix the palette issue with dark/light images - https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
The sync/adjust is available via right-click in the list view in export images (for image based subs only)
@von Suppé: Hm, not sure when your attachment will be approved... also tried to add the "sort by actor" in above beta.
@tormento: Only some subs or all subs?
Music Fan
12th October 2017, 10:22
@Music Fan: I've tried to fix the palette issue with dark/light images - https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
The sync/adjust is available via right-click in the list view in export images (for image based subs only)
Great, everything seems ok now :thanks:
By the way, you forgot to add the sup resize as a new option in the changelog of v 3.5.3.
And you can add too the delay for sup in the next changelog ;)
hello_hello
13th October 2017, 01:15
Thanks for the new version.
There's possibly a very minor bug (applies to previous versions too) when choosing a colour for ass subtitle styles.
Open the window for creating/editing ass styles. Click on something that opens the colour chooser.
There's a field that displays the hex value of the colour. If you right click on it you can copy the hex value but the paste option doesn't work. Is that by design? The paste option isn't greyed out when right clicking which indicates it should work.
Thanks again.
von Suppé
13th October 2017, 07:03
Hm, not sure when your attachment will be approved...
Yes, strange indeed. It's just the four lines .srt file I created myself with harmless text. Maybe the the mods have overlooked or are very busy?
Anyways, you can read it's content in the first picture; it's a screenshot from SE's listview from the same .srt.
Edit: BTW click on the picture to get full image.
Will go check your newest beta, thank you Nikse.
sfatula
14th October 2017, 08:27
I don't see any references to this in the thread, but, just an FYI in case someone is interested. SE 3.5.4 works fine via Wine, only had to install dotnet40 via winetricks. Spell check and dictionary and ocr all works.
tormento
21st October 2017, 12:27
@tormento: Only some subs or all subs?
All srt
Nikse555
22nd October 2017, 21:19
@tormento: Could you post a screenshot or the exact text (ctrl+c if you get a msgbox)?
mood
23rd October 2017, 00:03
@Nikse555
In Multiple replace window, what is the difference click on ok or apply button???
added on git-2c3ef90
I don't see any difference, why we have now 2 button for the same thing? ;)
Why apply the modifications to all uncheck box???
johner23
23rd October 2017, 04:58
Hi, dear all.
@Nikse555
Subtitle Edit can handle ( now or in future ) with UHD Blu-Ray subtitle format?
Any plans to add support ( if needed or desired ) to such subtitle format into Subtitle Edit engine?
Thanks.
Best regards.
devil (johner)
Nikse555
23rd October 2017, 15:58
@mood: The "Apply" button is useful if you have many rules and want to apply in a special order - or just do some at a time (and not all).
Unchecked items are now not applied - thx :)
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
@johner23: Is that a special format? If yes, is a sample available somewhere?
johner23
23rd October 2017, 22:47
Hi, dear all.
@Nikse555
Is that a special format? If yes, is a sample available somewhere?
I don't know yet, because I do not have specific hardware to open such discs.
But now that exist one program than can handle with some specific discs ( a "generic" crack for AACS 2.0 is not available yet I guess ), maybe someone could get interest and try to rip subtitles from such discs.
Maybe someone in the links below that have such hardware and have ripped all disc content could provide some subtitles samples for you.
I really don't know if "normal" blu-ray subtitles and UHD Blu-Ray subtitles are similar or not. If you find some sample you can analyse it better and see if Subtitle Edit can handle with them or not.
---> https://forum.doom9.org/showthread.php?t=174574
---> https://forum.videohelp.com/threads/385273-Russian-company-claims-that-they-have-software-to-decrypt-UHD-blu-rays
Thanks.
Best regards.
devil (johner)
MounaLafi
27th October 2017, 14:58
Hi,
I know that a lot of people who are interested in converting Arabic Subtitles (idx/sub) to text using OCR in SE through Tesseract, are diffidently familiar with the accuracy issues in Tesseract V. 3.x
Recently, I've encountered a person in Subscene who is claiming to successfully incorporate Tesseract V. 4.x within SE.
Tesseract V. 4.x has great improvement in regard of Arabic OCR accuracy.
I've made several attempts myself to incorporate Tesseract V. 4.x within SE, but in vain.
The issue is that the guy who did it is requesting money to teach the way of doing it and also to provide training files for each font.
You can see the videos here:
https://www.youtube.com/watch?v=NRh5njbC_V4
https://www.youtube.com/watch?v=wz8akEYraZw
https://www.youtube.com/watch?v=QEpbaKNEp0o
https://www.youtube.com/watch?v=mIffR8NOpxU
I would appreciate if someone know how incorporate Tesseract V. 4.x within SE and willing to share his way with us.
Thanks & Regards,
Mouna.
Nikse555
31st October 2017, 11:58
@MounaLafi: I've uploaded a test version here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBetaT4.zip (but... no support on this as I do not really want to look at alpha version problems!)
EDIT: You basically take the Tesseract.exe + dictionaries from beta 4. If you take the dictionaries from best ( https://github.com/tesseract-ocr/tessdata_best ) you must rename them to three letters as expected by SE.
tormento
31st October 2017, 20:19
@tormento: Could you post a screenshot or the exact text (ctrl+c if you get a msgbox)?
@MounaLafi: I've uploaded a test version here
My exception window disappeared with this version.
Will test more in the next days.
MounaLafi
1st November 2017, 18:44
@MounaLafi: I've uploaded a test version here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBetaT4.zip (but... no support on this as I do not really want to look at alpha version problems!)
EDIT: You basically take the Tesseract.exe + dictionaries from beta 4. If you take the dictionaries from best ( https://github.com/tesseract-ocr/tessdata_best ) you must rename them to three letters as expected by SE.
Thanks for your prompt response Niske.
I've attached a comparison between SE Beta Version and SE 3.5.4 below.
It is clear to me (as expected) the lines that has been identified and OCR'd are almost 100% accurate.
I just want to ask about the lines that were not OCR'd at all and about the lines that are giving weird results like:
EEE ESE
CIN أصبح? BEI
CT O?
RENE?
I understand that it is up to me now to enhance the accuracy of OCR by improving the Training File, but I just want to understand the cause of those errors as they were not occurring in SE 3.5.4
https://image.ibb.co/nh9ZJG/SE_Beta.jpg
https://image.ibb.co/dG9OCb/SE_3_5_4.jpg
macrea
7th November 2017, 18:17
I'm trying to "import/OCR subtitles from VOB/IFO (DVD)" but when I select the IFO file I'm given 2 framerate choices for importing - PAL (25fps) and NTSC (29.97fps). My DVD is 23.976fps. What should I do?
EDIT: never mind - I've got it figured out. It should be set to NTSC.
locotus
7th November 2017, 18:38
Hi, other thing is that on: File>export>vobsub/idx, program is not
remembering the value user set for "Bottom margin"
Thanks.
Lucius Snow
8th November 2017, 00:05
Hello Nikse555,
When rendering Blu-Ray SUP or Final Cut Pro PNG, a subtitle with only one line will display at the lower one of the two lines. Is it possible to change this and make it appear centered between the first and the second line?
i.e. current with 2 lines :
--- Subtitle with 2 lines
---
--- Subtitle with 2 lines
i.e. current with 1 line :
--- [Emty]
---
--- Subtitle with 1 line
i.e. wish with 1 line :
--- [Emty]
--- Subtitle with 1 line
--- [Emty]
I hope you understand what I mean.
Thanks.
tormento
24th November 2017, 15:33
If you remember, I told you before I had a crash with a srt file.
Now I can provide you a working example (https://ufile.io/jrgz5).
I can reproduce on 3.5.4 build 61.
open file
CTRL + SHIFT + H
APPLY
CTRL + SHIFT + F
APPLY
BOOM!
Same file works perfectly on 3.5.3 build 131.
Nikse555
24th November 2017, 19:31
@tormento: thx for the file :)
Crash should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
tormento
26th November 2017, 14:17
@tormento: thx for the file :)
Crash should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
It is. Will keep you informed if any crash will happen again.
Matt Kirby
22nd December 2017, 16:12
I've got a problem with the tool function "split long lines"
It splits where it shouldn't.
I have this subtitle item for example
- Niemals. -Ich bin aus einem
einfachen Grund nicht mit Euch ...
look at the photo
each line has less than 40 chars, but SE splits it after "-Niemals." and removes the "-" too.
When I remove the simple dot, after "niemals" It works fine, because it doesn't split this item. Is this a bug of SE or I don't understand the logic behind the split-function?
Ghitulescu
22nd December 2017, 17:22
- Niemals. -Ich bin aus einem
einfachen Grund nicht mit Euch ...
each line has less than 40 chars, but SE splits it after "-Niemals." and removes the "-" too.
This is a feature not a bug and it has been asked for many times in the past, see this thread....
Matt Kirby
22nd December 2017, 17:42
Then it's useless for me...
Boulder
22nd December 2017, 17:57
You are much better off creating a ticket at https://github.com/SubtitleEdit/subtitleedit, it's more likely to get a solution there.
Ghitulescu
22nd December 2017, 17:58
I do not use it either...
Generally I do by hand those hundreds of text lines, no "beautifier" ever worked for me.
Nikse555
22nd December 2017, 21:20
This one subtitle line:
- Niemals. -Ich bin aus einem
einfachen Grund nicht mit Euch ...
SE splits to two lines:
Niemals.
Ich bin aus einem einfachen
Grund nicht mit Euch ...
"-" is very common a start of a dialog where two different persons talk in the same subtitle line, like this:
- Hi Joe!
- Hi Jane!
So if this line is split to two, the "-" should be removed as only one person speaks in each line.
@Matt: I hope that explains the logic behind the split-function.
I'm not sure how the "-" is used in your subtitle...
Matt Kirby
23rd December 2017, 03:05
Yes, it's used for a dialog here too, but I like it if another person starts that he get's his own "-" so I can see there's now the other person speaking.
But the other question is why does SE split it anyway? My split value is "max 40 chars per line". (and 90 for both lines) Here it is less and it splits.
Thunderbolt8
23rd December 2017, 14:36
this case
- Niemals. -Ich bin aus einem
einfachen Grund nicht mit Euch ...
should be ideally split like that:
- Niemals.
- Ich bin aus einem einfachen Grund nicht mit Euch ...
simply because the hyphen indicates that another person is speaking and the speech pattern for a different person belongs into a different line, for the sake of clarity of comprehension.
johner23
30th December 2017, 03:34
Hi, dear all.
I have captured a video stream, and saved the subtitles in a text file.
The subtitle stream is labelled Type 'html' and Cause 'xhr'.
The first three entries are:
<a href='#' begin="5.706" end="8.289"><span class='ts'>00:05</span> (music)</a>
<a href='#' begin="13.037" end="14.961"><span class='ts'>00:13</span> - A Swiss scientist</a>
<a href='#' begin="14.961" end="17.128"><span class='ts'>00:14</span> had a marvelous statement,</a>
Can anyone tell me how to convert this format to srt? Subtitle Edit can do it properly?
Thanks for your help.
Best regards.
devil (johner)
Nikse555
30th December 2017, 13:17
@johner23: OK, try latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
It could be helpful to see a complete file...
johner23
31st December 2017, 06:30
Hi, dear all.
@ Nikse555
Thanks for your nice help! See the complete sample for carefully testing here (https://files.videohelp.com/u/191845/test.txt)
Best regards.
devil (johner)
von Suppé
16th January 2018, 13:17
Hi nikse,
I'm having trouble importing ASS styles which I earlier created with SE.
Is it possible that SE, when working with *.ass files, will auto-load evry ASS style I created? Must they be in a specific folder before start-up or something?
Much obliged,
cheers
sneaker_ger
16th January 2018, 13:29
For me ASS style export doesn't work at all (latest beta). It creates an ASS file but the line for the style is missing.
Workaround:
Just open actual ASS subtitle file with the styles you want to use and delete all lines (or import on it).
Apart from not working at all the export function is tedious because it seems you can only select one style. Same for import. ASS files often have many styles.
Nikse555
16th January 2018, 18:42
@von Suppé/sneaker_ger: I've updated the beta, to handle multi import/export - https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
How does that work?
Also, is that bug regarding not saving new styles fixed sneaker_ger?
sneaker_ger
16th January 2018, 19:02
Partly. Exporting styles that have spaces in their name doesn't work.
Nikse555
16th January 2018, 19:46
@sneaker_ger: thx for testing :)
Beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
sneaker_ger
16th January 2018, 19:53
Looking good. Thx.
von Suppé
17th January 2018, 14:00
@von Suppé/sneaker_ger: I've updated the beta, to handle multi import/export - https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
How does that work? Running a few first fast tests, I can confirm it works like a charm. Thank you, will test further of course.
von Suppé
17th January 2018, 14:42
The issue I mentioned earlier here https://forum.doom9.org/showthread.php?p=1820856#post1820856 still exists in latest beta.
Now, the srt file attachments in mentioned post still await approval(??) but thing is that, when wanting to create xml/png I first have to export as SUP, thereafter convert with BDSup2Sub.
As said, this tool (BDSup2Sub) has a sort-of frame-timings issue/bug, in that sense, that you have to specifically check the "Change frame rate" box and state both FPS Source and FPS Target rates, if a framerate other than 23.976 is used in both source and target fps.
If you are using BDSup2Sub internally in SE for xml/png export, can you "force tell" the tool to use this FPS Source/Target setting?
Nikse555
18th January 2018, 22:39
@von Suppé: SE does not use BDSup2Sub internally for xml/png export. SE (nearly) always write in "real time" and not "media time/drop-frame" - but I've made a test version here with a "Use media time (df)" check box: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
How does that work?
von Suppé
19th January 2018, 13:18
...I've made a test version here with a "Use media time (df)" check box: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.4/SubtitleEditBeta.zip
How does that work?
Will go test.
Thank you for fast respons :-)
Edit: First quick tests show XML/PNG output is still faulty, whether I choose df or not...
Zetti
27th January 2018, 16:53
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.5
von Suppé
31st January 2018, 11:45
Hi Nikse 555
When I preview subtitles for 4K video (4K video opened in SE, settings --> videoplayer 'mpv' with 'mpv handles preview text' checked) the subtitles don't scale in proper ratio to the video. They seem twice as big.
With 1080 I don't have this issue. Can you do something about that? You must know that I am previewing on a 1920x1080 monitor.
von Suppé
31st January 2018, 13:31
Hi Nikse555,
Another thing I noticed. While your new import & export for ASS styles work like a charm now, I encounter this issue.
For example,
If I create a style in Aegsub, I can set the bottom offset to let's say 288. I want this for 4K video with a 2.40 movie aspect ratio, so the subtitles are shown above the black bar (which itself is 280 pixels in this case, I add 8 to lift it slightly extra).
This ASS style I can directly import in SE and it reacts well on this offset. I can see that in the SUP export window, where - eventhough "Bottom margin" is greyed out - it says 288. And off course, checking the output in BDSup2Sub, the result is okay.
Creating an ASS style in SE, in the style properties window, the bottom offset is represented by the box "Margin vertical". However, I can't get the value higher than 250.
Is there a specific reason why you implemented this maximum and is it possible to take it away, or at least increase that value?
Nikse555
3rd February 2018, 12:15
@von Suppé: I've upped the margin max size to 500, is that enough?
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.5/SubtitleEditBeta.zip
I don't know about the preview - it should just use the "Subtitle preview font size" from Options -> Settings -> Video player.
von Suppé
4th February 2018, 11:07
That should be enough.
Thanks Nikse555 :)
von Suppé
6th February 2018, 13:55
The margin max size is enough, it works. Thanks.
I think I found a bug.
During editing of a *.ass file, when applying 'Tools --> Dialogues - remove hyphen in first line of dialogues...' (it's a plugin) all styles are vanished from the styles list.
Nikse555
6th February 2018, 21:35
During editing of a *.ass file, when applying 'Tools --> Dialogues - remove hyphen in first line of dialogues...' (it's a plugin) all styles are vanished from the styles list.
Thx :)
Plugin updated. Go to File -> Plugins to update.
varekai
7th February 2018, 14:56
@ Nikse555
Hello!
Thanks for Subtitle Edit, it's one-of-a-kind!
Really enjoy making my own subtitles for my projects.
One thing I've been thinking about is the space below the second subtitle row.
It cuts the bottom of the subtitle and it would be great if it could have the same space as the first row have.
Maybe there are some settings for this but I can't find any.
Here's a screenshot:
https://frupic.frubar.net/shots/36437.png
Best regards
von Suppé
7th February 2018, 15:45
Thanks Nikse555 :)
@ varekai,
When I see your picture I think it is because you didn't choose your player to be full screen. Or the interface is over the subs. I highly doubt if SE will chop off the lower part of any sub as shown.
You can also tell your player to show the subs a bit higher in settings.
varekai
7th February 2018, 16:23
Thanks Nikse555 :)
@ varekai,
When I see your picture I think it is because you didn't choose your player to be full screen. Or the interface is over the subs. I highly doubt if SE will chop off the lower part of any sub as shown.
You can also tell your player to show the subs a bit higher in settings.
Hmm... I don't understand what you mean?
Here's what I do:
Open subtitle (srt)
Video --> Open video file (m2ts)
If I chose "Un-dock video controls" I get the same result.
DirectShow is the "player" and I don't know how to change "player" to full screen.
It looks like the "player" window/bar doesn't adapt to 2 subtitle rows?
Screenshot of single subtitle row:
https://frupic.frubar.net/shots/36438.png
Edit: Well... guess I'll have to set "Subtitle preview font size" to 17 to fit, although the small size does hurt my eyes...
Wish it could be set to bigger size and still keep it viewable inside DirectShow window.
Regards
von Suppé
8th February 2018, 12:58
@ varekai
Aahh I see now. It's the preview of SE's chosen videoplayer itself. Sorry for the misinterpretation.
What you can do is in settings, choose mpv as videoplayer (you have to click download button). And then also set "mpv handles preview text.
MPV will show all subtitles above the interface.
cheers
varekai
8th February 2018, 13:38
@ varekai
Aahh I see now. It's the preview of SE's chosen videoplayer itself. Sorry for the misinterpretation.
What you can do is in settings, choose mpv as videoplayer (you have to click download button). And then also set "mpv handles preview text.
MPV will show all subtitles above the interface.
cheers
Thanks, tried mpv before and subs preview are not so elegant as with DirectShow. Guess I can try to get used to it once more...
von Suppé
8th February 2018, 16:36
I know what you mean. Using SE for creating/editing srt's, I don't mind much how it looks in my preview because the end goal is to play it with mediaplayers.
And as for srt, which is only text and timing-info, I always end up creating SUP files to mux with video. That way I control how & where the subtitles will look.
Maybe there is some settings withtin the DirectShow configuration you can change for how subtitles are displayed?
varekai
8th February 2018, 17:02
@ von Suppé
Yeah, an improved DirectShow subtitle preview would be great, or a better looking mpv preview that is smoother for my eyes.
When editing and syncing it gets tiresome after a while for my eyes, so the smoother the preview font style is the better.
So I'll stick to DirectShow for now.
I do much the same as you, starting with either an imported sup/sub/ or from scratch an srt.
When needed I create a BD sup where I can choose font, style, size, placement, color etc.
Without the excellent Subtitle Edit I would be lost! :D
Zetti
27th February 2018, 20:13
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.6
tormento
28th February 2018, 17:43
@Nikse555
Could you please enable CTRL+a as "select all" in windows such as batch convert? Would be really comfortable to use when removing lot of subs.
There is a regression, dunno since which version. If you feed the following srt:
1
00:01:48,275 --> 00:01:51,987
And why would I want them? DRIVER:
They were trying to grab your prize.
2
00:18:04,459 --> 00:18:06,086
What? BLAKE: Name's Jimmy.
3
00:21:21,572 --> 00:21:22,949
Go! SWAT 2: Go!
4
00:31:45,738 --> 00:31:48,324
Now. YUPPIE: Creep!
5
01:03:23,258 --> 01:03:26,470
Mr. Wayne, over here! MAN 2:
How's it feel to be one of the people?
6
02:03:45,377 --> 02:03:46,753
Bruce. WAYNE: You okay?
and you try to remove text before colon (":"), if you select "Only if text is UPPERCASE, SE won't find any. If you deselect it, SE won't find all and the ones it finds give wrong output.
raymondjpg
11th March 2018, 06:14
I've noticed that on a number of occasions in this forum the question has been asked can an option be implemented in the drop-down menu for dictionaries, or in the settings, to return to or save the last used dictionary. At the moment the English US Dictionary is always loaded by default.
I haven't been able to find any response to this question. Would it be too difficult to implement?
hello_hello
13th March 2018, 15:01
I get this error message when running Subtitle Edit 3.5.6 on XP, although once I dismiss the error message, it seems to run normally (so far). I'm using the portable version.
Thank you.
https://s9.postimg.org/vh9jevn5b/Subtitle_Edit_Exception.gif
hello_hello
13th March 2018, 15:24
I've noticed that on a number of occasions in this forum the question has been asked can an option be implemented in the drop-down menu for dictionaries, or in the settings, to return to or save the last used dictionary. At the moment the English US Dictionary is always loaded by default.
If you're referring to the dictionary selected when opening the spell checker, I was one of the people who made a similar request, and it was changed quite a while ago.
https://forum.doom9.org/showthread.php?p=1765592
raymondjpg
14th March 2018, 03:30
If you're referring to the dictionary selected when opening the spell checker, I was one of the people who made a similar request, and it was changed quite a while ago.
https://forum.doom9.org/showthread.php?p=1765592
Thanks for the response but:
1. I cannot locate that beta from 2016 from your link
2. The latest version 3.5.6 still defaults to English US.
3. The latest beta also defaults to English US.
Perhaps I am just unable to find the setting to have Subtitle Edit open to last used dictionary. If so, could you please point me to it?
TIA
jpsdr
14th March 2018, 09:45
BTW, as we are now at version 3.5.6, shouldn't the thread title be updated ?
raymondjpg
14th March 2018, 12:32
If you're referring to the dictionary selected when opening the spell checker, I was one of the people who made a similar request, and it was changed quite a while ago.
https://forum.doom9.org/showthread.php?p=1765592
My mistake. Your request was regarding the dictionary selected when opening the spell checker.
My request is for the dictionary selected in the OCR dialogue window. Can that default to the last dictionary used there too?
Nikse555
14th March 2018, 15:43
@varekai: the video preview font for "mpv" can be modified by modifying the default ASS/SSA font in Options -> Settings, so maybe you can find a nicer font/font-size.
@tormento: sorry, SE never had this feature - I think I tried once, but it gave some false positives...
@hello_hello: The startup error is due to the fact, that TLS is not supported in .net framework 4.0 (and the newest .net framework available in WinXp). Next update will not start with this crash but some downloads from SE will no longer work - read more here: https://github.com/SubtitleEdit/subtitleedit/issues/2799
@raymondjpg: SE actually does save/use the last used dictionary. Unless the language is auto-detected, SE will use the last used dictionary.
So, when do you get the wrong language and what language do you use?
varekai
14th March 2018, 17:13
@varekai: the video preview font for "mpv" can be modified by modifying the default ASS/SSA font in Options -> Settings, so maybe you can find a nicer font/font-size.?
Thanks for the info, will try that with my next project.
Also thanks for an outstanding one-of-a-kind software!
Best regards
hello_hello
14th March 2018, 20:29
@raymondjpg: SE actually does save/use the last used dictionary. Unless the language is auto-detected, SE will use the last used dictionary.
So, when do you get the wrong language and what language do you use?
I'd forgotten, but he seems to be correct.
When importing an idx/sub subtitle, the OCR window always seems to default to using the English US dictionary (I tried a PAL disc).
I guess it's got to pick the correct language, but it would be nice if it remembered the last dictionary used for a particular language (assuming there's more than one), such as the English GB dictionary.
Cheers.
raymondjpg
15th March 2018, 01:51
@raymondjpg: SE actually does save/use the last used dictionary. Unless the language is auto-detected, SE will use the last used dictionary.
So, when do you get the wrong language and what language do you use?
I seem to get English US dictionary in the OCR dialogue whenever I open a .sub file with English language. Sometimes those .sub files are for UK English language, and I prefer to use the English UK dictionary.
It seems that the OCR dialogue in SE is correctly identifying the language, but if I select the English UK dictionary in the drop down menu, the next time I open the OCR dialogue with the same .sub file I get the English US dictionary back by default.
I'd forgotten, but he seems to be correct.
When importing an idx/sub subtitle, the OCR window always seems to default to using the English US dictionary (I tried a PAL disc).
I guess it's got to pick the correct language, but it would be nice if it remembered the last dictionary used for a particular language (assuming there's more than one), such as the English GB dictionary.
Cheers.
Thanks for the confirmation.
tormento
15th March 2018, 15:54
@tormento: sorry, SE never had this feature
Could you implement it? CTRL+a is a sort of universal shortcut in lot of programs.
I think I tried once, but it gave some false positives...
Just add "remove any CAPITAL TEXT followed by colon".
Nikse555
15th March 2018, 20:29
@raymondjpg: I've tried to remember spell check dictionary in ocr in latest beta.
@tormento: Ctrl+a (select all), ctrl+d (de-select) , ctrl+shift+i (invert) should now be implemented in many list views in latest beta.
Latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.6/SubtitleEditBeta.zip
raymondjpg
16th March 2018, 04:36
@raymondjpg: I've tried to remember spell check dictionary in ocr in latest beta.
Latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.6/SubtitleEditBeta.zip
Thank you! That looks like it is working now.
tormento
17th March 2018, 12:44
@tormento: Ctrl+a (select all), ctrl+d (de-select) , ctrl+shift+i (invert) should now be implemented in many list views in latest beta.
:thanks:
raymondjpg
23rd March 2018, 05:56
I am trying to reset Subtitle Edit 3.5.6 (the installed version) to default (initial installation) settings.
I uninstall Subtitle Edit (I have tried both from Control Panel in Windows 7 64 bit, and from the uninstaller in the Program Files Directory), using the option to clear all modifications and settings, and also delete the User directory. I then cleanup temporary files in the Windows 7 installation and reboot the PC.
When I then reinstall Subtitle Edit 3.5.6, the program remembers the last opened directory.
It appears that there is something of Subtitle Edit's settings remaining in the system after a complete uninstall. If anyone knows what it is, could they please point me to it.
TIA
varekai
23rd March 2018, 09:48
It appears that there is something of Subtitle Edit's settings remaining in the system after a complete uninstall. If anyone knows what it is, could they please point me to it.
TIA
Uninstall then delete this Subtitle Edit folder:
C:\Users\***(user name)***\AppData\Roaming\Subtitle Edit
raymondjpg
23rd March 2018, 12:08
Uninstall then delete this Subtitle Edit folder:
C:\Users\***(user name)***\AppData\Roaming\Subtitle Edit
Thanks for the response, but yes, I have removed that folder after uninstalling SE. The reinstalled SE still remembers the last opened directory.
varekai
25th March 2018, 12:24
Thanks for the response, but yes, I have removed that folder after uninstalling SE. The reinstalled SE still remembers the last opened directory.
Have you tried to locate and delete Subtitle Edit in regisrty?
HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Uninstall\SubtitleEdit_is1
raymondjpg
26th March 2018, 06:08
Have you tried to locate and delete Subtitle Edit in regisrty?
HKEY_LOCAL_MACHINE\SOFTWARE\Microsoft\Windows\CurrentVersion\Uninstall\SubtitleEdit_is1
Thanks, but that entry does not appear to be in the registry after an uninstall. It reappears after SE is reinstalled.
mkver
26th March 2018, 08:18
Have you just searched your registry for "SubtitleEdit"? This seems to be a Windows feature and not something that SE controls.
I remember that Windows keeps MRU (most recently used) lists in the registry.
raymondjpg
26th March 2018, 09:38
Have you just searched your registry for "SubtitleEdit"? This seems to be a Windows feature and not something that SE controls.
I remember that Windows keeps MRU (most recently used) lists in the registry.
That may be it, although I could not find such an entry in the registry specifically for SubtitleEdit after I had completely uninstalled the program.
sfatula
28th March 2018, 22:31
When using Tesseract on SE 3.5.6 under Wine on Mac OSX, I can add to user dictionary, it retains changes, etc. However, adding to names/noise list does not seem to update the file so this info is lost for future uses. Is it supposed to update the file thats lists names/noise words?
Nikse555
31st March 2018, 12:20
@raymondjpg: mkver is most likely correct - SE (except installer) does not save anything in the registry (all settings are saved in the Settings.xml file).
@sfatula: thx, that's a bug. "Add to names list" in OCR spell check did not work - should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.6/SubtitleEditBeta.zip
x265
22nd July 2018, 03:37
How do i convert .srt file to dvd .sup in subtitle edit? I want to add subtitles to a DVD. I tried exporting as bd sup but that does not load correctly in muxman.
varekai
22nd July 2018, 12:44
How do i convert .srt file to dvd .sup in subtitle edit? I want to add subtitles to a DVD. I tried exporting as bd sup but that does not load correctly in muxman.
Have you tried export to VobSub (sub/idx)?
x265
22nd July 2018, 14:32
Have you tried export to VobSub (sub/idx)?
Yeah i tried that but it won't open in muxman. It only accepts dvd .sup.
Music Fan
22nd July 2018, 18:45
Try export in Blu-ray sup, but I don't know if it's the same format than dvd sup.
x265
22nd July 2018, 20:33
Try export in Blu-ray sup, but I don't know if it's the same format than dvd sup.
Error message: subpicture type not supported.
Music Fan
22nd July 2018, 23:19
Then try an export in VobSub or Blu-ray sup and convert to Dvd sup with another tool (BDSup2Sub for example).
Ghitulescu
26th July 2018, 09:15
Try export in Blu-ray sup, but I don't know if it's the same format than dvd sup.
It's not.
Just why one would try to export in BD format when the target is DVD? What kind of logic in thinking is this?
Music Fan
26th July 2018, 10:25
I said : "I don't know if it's the same format", could I be more cautious ?
There are some common points between Dvd and BD, and I was hoping for him that a BD sup in 720x576 (or 720x480) could be dvd compliant.
Also, I made this suggestion thinking that Nikse555 couldn't have forgotten one of the most common formats (dvd sup doesn't appear in the list), thus I tought he had maybe given the same name to both formats and that only higher resolutions wouldn't be dvd compliant.
That's a very logical reasoning actually.
tormento
3rd August 2018, 15:37
@Nikse555
Found another srt that can crash Subtitle. Tried latest beta:
h**ps://openload.co/f/twwBdcwStS0/crash.zip
Just try to remove hearing impaired text
varekai
4th August 2018, 10:20
@Nikse555
Found another srt that can crash Subtitle. Tried latest beta:
Just try to remove hearing impaired text
No problem for me using SE version 3.5.6
You really should use another cloud service!
That openload site is a piece of s***!
Sorry for being a bit harsh...
You could get a free account on https://www.mediafire.com/upgrade/
There are some ads but not so pesky as on openload.
http://www.mediafire.com/file/fnxfiaerhb113lz/no%20crash.zip
Best regards
Nikse555
7th August 2018, 20:34
@tormento: thx for the info (crashed with some combination of options in 'remove text before colon').
Should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.6/SubtitleEditBeta.zip
tormento
9th August 2018, 13:53
You really should use another cloud service!
That openload site is a piece of s***!
OT: uBlock Origin is the cure.
tormento
9th August 2018, 13:54
@tormento: thx for the info
Thanks for fast fix. Wil try for some days before reporting.
P.S: I suggest you to add text removal before ":" also for mixed case option, such as "BiLBO:" and so on.
varekai
10th August 2018, 10:58
OT: uBlock Origin is the cure.
Looked into it and the uBlock Origin reviews aren't hot...
Why you won't use mediafire is beyond me.
Regards
rco133
13th August 2018, 07:36
Hi.
Got a question about the spell checking.
After OCR I often end up with a lot of lines with "I" (capital I) actually being a l (low capital l) or L (capital L).
Words like
l'm should be I'm
L'm should be I'm
ls should be Is
lt's should be It's
When running a spell check, it happilly accepts all those Words, eventhough none of them I think is even valid English Words.
Also "Bob." became "Bob-ll
Which might be valid, but should still be pointed out by the spell checker?
I am not even sure if spell checking is an internal part of Subtitle Edit. If not I guess there is nothing to do about it.
I just think that the spell checker accepts a lot of Words, that are not even valid English words.
Maybe it is me not using it correctly. If so please let me know how to make it catch all these words.
If not, can something maybe be changed to make it catch all these words?
Thanks in advance.
rco133
Nikse555
13th August 2018, 18:20
@rco133: SE uses Hunspell... I would prefer not to add extra code with special OCR related hardcoded rules in the spell check.
ATM you can either:
1) Use Tools -> Fix common errors - Fix common OCR errors
or
2) Make your own rules via Edit -> Multiple replace
Ghitulescu
14th August 2018, 07:57
Probably half of this thread is concerned with partial solutions to a problem that should never existed in the first place.
Why do these spelling errors arrive and wherefrom? Can't the user correctly write in his own language?
Of course I know the answer. The errors comes from OCR. Geez, we do not have enough problems, we are bored, let's convert from the original subtitles to another one, to save only a couple of kB, to occupy our minds then with issues like code page (for anyone not English), L or I, # or ♫, colours, timings, positions and of course the compatibility of the result with the existing players. Players that otherwise will play the original file just perfect, and 83-years-old-age compliant.
Some people want the dash before the first line of dialogue, others request its deletion. Go figure.
Issues and problems that concern a secondary if not even non-related task of the software, a bonus functionality if you like.
Meanwhile issues that are part of the core, like loosing the position on-screen of the original subtitles (these are centered after processing) do not have time to get solved, because someone has a spelling issue with tembo-tembe subdialect and semicolons :) .
DMD
26th August 2018, 14:01
Good morning.
I have a problem with DirectShow (W10 @ 64bit) since a few days has stopped working and in the warning of Subtitle Edit I get a black screen but the audio is present.
I am forced to set up VLC as a video player in the Options> Preferences.
I can not understand what happened, :( can someone give me some advice?
Thank you
Nikse555
26th August 2018, 14:51
@DMD: It's probably because you're missing a decoder for the video - LAV filters contains decoders for the most common used formats - get it here: https://github.com/Nevcairiel/LAVFilters/releases (use the installer)
(if you're sync'ing or creating subtitles I recommend mpv as VLC does not have precise seeking)
DMD
26th August 2018, 14:55
Thank you for your help.
I installed LAV filter 0.72, but I still have the same problem.
https://s20.postimg.cc/hyf3lhbzh/Screenshot_001.png
https://s20.postimg.cc/a5oftidq5/Screenshot_002.png
Nikse555
26th August 2018, 15:09
What version of LAV filters did you install? SE probably runs as a 64-bit program on your computer, so it's important you did not install only the 32-bit version. (I think the installer installs both versions per default).
MediaInfo can tell you what video codec you're having trouble with.
Perhaps someone else has some good ideas?
LowDead
26th August 2018, 22:44
No good ideas, but I actually have the same problem since a while back.. Haven't gotten around troubleshooting yet though.
//LD
DMD
27th August 2018, 13:26
FINALLY I RESOLVED !! :):)
After endless tests, I discovered that the problem was due to the latest nVidia driver release and not a codec problem.
The latest drivers are version 398.82 for the GeForce GTX 960 video card, I tried to install version 390.65, with this version I solved all the problems.
Thank you all for the availability
varekai
28th August 2018, 09:12
The latest drivers are version 398.82 for the GeForce GTX 960 video card,
I tried to install version 390.65, with this version I solved all the problems.
Not the drivers fault, somethings wrong with your setup.
Works great for me with 398.82.
BTW, latest drivers are 399.07.
Update:
Works perfectly with 399.07.
Regards
DMD
29th August 2018, 13:46
I created the current image of the pc through Acronis, I have formatted and reinstalled only the 399.07 Driver the video with directShow turns black.
With the regressive driver test, I checked the regular operation with version 391.35.
I think the problem could be the GTX960 video card that has problems with newer drivers.
Regards
Zetti
9th September 2018, 23:13
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.7
jpsdr
10th September 2018, 12:43
It never happened to me before, i've until now always used the portable version, but with this last version, i can't start OCR, i have popup and error message telling me that tesseract can't be started, to check everything is properly installed.
Everything is properly installed (redistribuable 2017, Framework .NET 4.7.1). Under Windows 7 x86.
Rolled back to 3.5.6, everything worked fine again.
Edit :
Could this be related ?
Top Issues Fixed in 15.8.3
These are the customer-reported issues addressed in 15.8.3:
Visual Studio 2017 version 15.8.2 contained a pre-release build of .NET Core SDK 2.1.401 that is incompatible with Visual Studio. We have corrected this issue with Visual Studio 2017 version 15.8.3.
varekai
10th September 2018, 15:49
Many thanks for the new release! SubTitle Edit is one-of-a-kind!
Installed 3.5.7 and kept old settings.
Started a new project today and noticed some new stuff in GUI.
Didn't change anything.
Ran the new project and there was a noticeable sluggish performance.
Usually it just flows through but now it stumbled and almost froze for a second before moving on.
Tested with old projects and the sluggish behavior is there.
So I rolled back to 3.5.6 and all is good, both the new project and old ones and the sluggish behavior is gone.
I have no clue on what this is about or what I can do.
Best regards
Nikse555
10th September 2018, 16:39
@jpsdr: The Tesseract version used by SE is now "Tesseract 4.0 beta 3" which might need a newer "Visual C++ Redistributable for Visual Studio" - could you try to install this: https://www.microsoft.com/en-us/download/details.aspx?id=48145
Does that help?
The reason for upgrading from Tesseract 3.02 to 4 is mainly due to support for more languages... Tesseract 4 is a lot slower and it's also the reason for the SE increase size (4 extra mb)!
@varekai:The sluggish behavior is when doing what exactly? Did you change video player?
varekai
10th September 2018, 16:48
Sorry for bad explanation...
I have first window open (List view) and then File --> Import/OCR Blu-ray (.sup) subtitle file...
No video player active in this scenario, when I use video I have mpv selected.
Havn't tried that yet in new version but it worked perfect in 3.5.6 last time I used it.
jpsdr
10th September 2018, 16:52
Euh... Your link is for a VS2015 redistribuable. You can't install redistribuable 2015 when you've installed the redistribuable 2017. I have installed the 14.15.26706 version of 2017 redistribuable, i have Framework.NET 4.7.2, tesseract still don't want to start.
Not a big deal, for now, i'll stay with 3.5.6.
jpsdr
11th September 2018, 10:13
More information : I have a big popup with a lot of informations, but the first line, translated said : This is not a valid application for this operating system.
As i've notify in post #716, i am under Windows 7 x86.
Meaning your tesseract version either is not compatible Windows 7, or is not for x86 OS. Have you tested either under Windows 7 and/or under 32 bits OS ?
After, if for now versions are not Windows 7 compatible or are not 32 bits possible anymore, Ok, but at least tell it.
Nikse555
12th September 2018, 06:17
@jpsdr: thx, it seems that I've included the x64-bit version instead of the x86 version of Tesseract... I'll upload a new beta later today.
varekai
12th September 2018, 12:45
@Nikse555
The reason for upgrading from Tesseract 3.02 to 4 is mainly due to support for more languages...
Tesseract 4 is a lot slower and it's also the reason for the SE increase size (4 extra mb)!
Is this the reason I'm seeing a "sluggish" performance?
Are there any settings in new version I could use have the flow as in 3.5.6?
Best regards
Nikse555
12th September 2018, 14:57
@jpsdr: Could you test latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.7/SubtitleEditBeta.zip
How is that?
@varekai: You could try the above beta.
varekai
12th September 2018, 16:20
@varekai: You could try the above beta.
Hello!
Just tried the beta and experienced the same sluggishness.
Imported a Blu-ray sup.
Also I noticed a lot of weird linebrakes that is not in 3.5.6
This is a screenshot from 3.5.7
https://i.imgur.com/Q4sY5Ek.jpg
This is from 3.5.6
https://i.imgur.com/yRJ5mCK.jpg
Best regards
DMD
12th September 2018, 18:13
Good morning.
I'm sorry, for the trivial question, but I can not select MPV player.
I downloaded the file from the site https://mpv.io/installation/
https://s20.postimg.cc/y7vpoo0j1/Screenshot_001.png
The version is in portable format, I also tried to press "Download mpv lib" but I can not select "mpv".
How should I proceed?
Thank you
Nikse555
13th September 2018, 04:35
@varekai: Can you upload the source image? (you can right click on it in SE and save it)
@DMD: You need to click "Download" again in the popup window.
You can also download the "mpv-1.dll" from the dev package - be sure to take the release version in the correct x64 or x86 version. Atm latest dev package is here: https://mpv.srsfckn.biz/mpv-dev-20180731.7z
DMD
13th September 2018, 07:41
@DMD: You need to click "Download" again in the popup window.
You can also download the "mpv-1.dll" from the dev package - be sure to take the release version in the correct x64 or x86 version. Atm latest dev package is here: https://mpv.srsfckn.biz/mpv-dev-20180731.7z
Thank you very much, I finally solved! :):)
jpsdr
13th September 2018, 09:04
Ok, i've tested, it works now, but my feedback on a sup file i've tested :
- Too slow.
- Too much more errors. Tesserac 3 produces realy significant better results.
For me, 3.5.7 is only negative. So, question :
- Is it possible to still use the previous tesserac version if i replace the files in the tesserac4 directory ?
- Is it possible to have both versions of tesserac, with two directory "tesserac" and "tesserac4", and a global setting in the "configuration/tools" which select the tesserac version used ?
varekai
13th September 2018, 14:00
@varekai: Can you upload the source image? (you can right click on it in SE and save it)
This is from 3.5.7
https://i.imgur.com/ZMtUhOK.png
Best regards
Nikse555
19th September 2018, 21:08
@jpsdr: OK, I've tried to add Tesseract 3.02 back in also... how does this work: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.7/SubtitleEditBeta.zip ?
jpsdr
20th September 2018, 09:43
When i select 3.02 in this beta version, i have more errors than with the 3.02 from the 3.5.6.
I've replaced the whole 3.02 (downloaded) directory by the directory from the 3.5.6 (for testing, just in case), but no change in behavior.
RyaNJ
20th September 2018, 12:01
Hi there.
I've got a quick question for you all. I've run into an issue with the BD subtitle import (.sup). When they are subtitles at the top and bottom of an image it doesn't seem to correctly identify this and either ignores one or puts them both into a single line at the bottom, leading to knock on issues. It could be that I'm doing something wrong since I am new to this tool. Does anyone have any experience with this? Is there a way around this that doesn't require manually editing these every time?
Thank you kindly!
Ghitulescu
20th September 2018, 17:24
I've run into an issue with the BD subtitle import (.sup). When they are subtitles at the top and bottom of an image it doesn't seem to correctly identify this and either ignores one or puts them both into a single line at the bottom, leading to knock on issues. It could be that I'm doing something wrong since I am new to this tool.
It appears that the spatial position of the subtitles is lost during importing.
RyaNJ
20th September 2018, 17:26
It appears that the spatial position of the subtitles is lost during importing.
That's more than slightly annoying. If there isn't any way around that then does anyone have a suggestion as to a tool that can actually do it?
This is by far one of the best editing tools that I've used so hopefully there is some way to make this work though :)
Ghitulescu
21st September 2018, 07:12
If there is one I haven't found it.
Nikse555
23rd September 2018, 09:26
@jspdr: Yes, SE 3.5.7 should probably not have included Tesseract 4 by default... I've uploaded a new beta here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.7/SubtitleEditBeta.zip
Do you still have more errors with 3.02 in the new beta than with the 3.02 from the 3.5.6? If yes, should you share your sub?
Other changes: If you right-click in the list view you can also check the new "image pre-procissing" + Tesseract should also be faster if you have a nice CPU with many cores.
@varekai: Yes, Tesseract 4 is not really super... you can change the engine mode to "Neural Nets LSTM" which is the new engine - but it does not have support for italics and it also seems to have problems with line breaking.
@RyaNJ/Ghitulescu: Sorry, SE does not support positioning of subtitles much. Do you need only like "align top, left" or the actual position like "100, 300" ?
Also, you could try https://www.videohelp.com/software/SubExtractor
RyaNJ
23rd September 2018, 09:28
@RyaNJ/Ghitulescu: Sorry, SE does not support positioning of subtitles much. Do you need only like "align top, left" or the actual position like "100, 300" ?
Also, you could try https://www.videohelp.com/software/SubExtractor
Thanks for the reply :) Nah, I don't need the exact position - I literally just need the subs at the top to stay at the top and the ones at the bottom to stay at the bottom. I could do it by hand but that would take days.
I'll give that a try, thanks!
varekai
23rd September 2018, 10:26
@varekai: Yes, Tesseract 4 is not really super...
you can change the engine mode to "Neural Nets LSTM" which is the new engine
- but it does not have support for italics and it also seems to have problems with line breaking.
Hello Nikse!
Thanks for the follow up and clarification, appreciate that!
For now I will stay with 3.5.6 which works very well for my needs.
Thanks and best regards
RyaNJ
23rd September 2018, 15:04
Thanks for the reply :) Nah, I don't need the exact position - I literally just need the subs at the top to stay at the top and the ones at the bottom to stay at the bottom. I could do it by hand but that would take days.
I'll give that a try, thanks!
I just played around with it a bit and while it works... well... let's say the resulting mess is worse than having to fix a few lines with simultaneous top/bottom subtitles here and there. It's OCR isn't the best (given it's age that isn't a surprise really) and the resulting mess is difficult to clean up.
I think I'll just have to give up on the idea and just leave the subtitles alone. I don't have a hour to spend cleaning up each one of these videos as that would take weeks. I just don't have that sort of time. If Subtitle Edit can ever handle detection of top/bottom subtitle elements then I will revisit in the future.
Thanks for the suggestion none the less :)
Matt Kirby
24th September 2018, 09:38
Is there a way to get "Auto break" splits the text in 3 lines when there are to many chars?
in options I set "max number of lines" to "3". Now I want that "auto break" uses this option, but it allways splits the text into 2 lines even if there are 50 chars per line
Three times 33 chars per line would be better for me.
DMD
24th September 2018, 13:49
Good morning.
I'm using version 3.5.7, when I load a ".sup" file to create the srt file via OCR, I get this message.
https://i.postimg.cc/SsDVVmYx/Screenshot_001.png
This does not happen with the previous version 3.5.6, how can you solve?
I noticed that in the installation folder, there is a new folder "Tesseract4" instead of "Tesseract", present in the previous version.
Thank you
https://i.postimg.cc/h40hyfGr/Senza_titulo-3.png
jpsdr
24th September 2018, 14:44
The last beta linked seems to behave like 3.5.6 with tesseract 3.02.
DMD
24th September 2018, 17:49
In version 3.5.7 I have set in the section
Engine mode> Default, based on what is available, it works regularly.
If I select Original Tesseract only (can detect italic) I get the error message.
mood
26th September 2018, 06:14
we need a fix version soon to remove Tesseract 4 and folders, to many bugs on Tesseract4
sneaker_ger
14th October 2018, 13:37
In the MPC-BE thread (https://forum.doom9.org/showthread.php?p=1854697#post1854697) an mp4 sample was posted. SubtitleEdit (latest beta) handles it differently than ffmpeg and I think ffmpeg is correct.
ffmpeg:
https://i.imgur.com/g8a2zEd.png
SubtitleEdit Beta:
https://i.imgur.com/FbjlpIZ.png
I'm not 100% sure the file isn't broken but maybe you want to take a look at it.
mkver
14th October 2018, 14:12
I vaguely remember that there is something odd with ffmpeg and timed text in mp4: The last subtitle would continue until the end of the file. So I would not count (or even bet) on ffmpeg being correct. But I was too lazy to open an issue for it.
sneaker_ger
14th October 2018, 14:14
I vaguely remember that there is something odd with ffmpeg and timed text in mp4: The last subtitle would continue until the end of the file. So I would not count (or even bet) on ffmpeg being correct. But I was too lazy to open an issue for it.
https://trac.ffmpeg.org/ticket/6341
?
But this problem is not specifically about the last line.
mkver
14th October 2018, 14:30
Yes, that seems to be it. I have already encountered files (which have been created by ffmpeg according to their metadata) whose last line continued for several minutes and for which ffmpeg and SubtitleEdit differed in their timings just like in your screenshots. Therefore they might be related somehow.
von Suppé
25th October 2018, 08:42
Hi Nikse,
It seems a few old bugs have sneaked in again in version 3.5.7.
The "Dialogue one hyphen only" plugin also removed the hyphen in a single line subtitle.
Also I encountered several times (not always, though) that, after running "Tools --> Fix common errors", the changes previously made by the "Dialogue one hyphen only" plugin are undone. So I would have to run the plugin again.
In version 3.5.6 I didn't run into these issues.
Can you please take a look?
Thanks
DMD
28th October 2018, 16:10
Good morning
I ask if it is possible to enable the font color panel to the hexadecimal code.
The possibility is to copy a color with photoshop in hexadecimal code and the possibility to insert it in the appropriate box.
What I currently do not allow, but only RGB values.
Thank you
Zetti
3rd December 2018, 01:12
Hello.
Both version 3.5.6 and 3.5.7 wont auto translate with Google Translate.
The lastest beta can only translate about 37 lines with Google Translate but without API Key.
And the Microsoft Translator wont working at all without API Key.
I have tried to get a free API Key to both Google Translate and Microsoft Translator, but it seems to hard for me.
Nikse555
4th December 2018, 19:18
@DMD: Where is this font dialog?
@Zetti: I've tried to squeeze a few more lines through (by using larger chunks) in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.7/SubtitleEditBeta.zip
How many lines do you get now?
Zetti
4th December 2018, 21:04
I getting a full subtitle now.
Edit: The subtitle have 2099 lines.
kenzofabio
9th December 2018, 07:20
For everybody in using subtitle edit 3.57 i dont know what happen what this program.i just can use one translate for subtitle.and it said that api quota acess.thx
Zetti
9th December 2018, 15:15
@kenzofabio
Check this github link: https://github.com/SubtitleEdit/subtitleedit/issues/3209
orion44
17th December 2018, 17:41
Thanks to the developer for this excellent program.
A feature request: When exporting subtitles to VobSub (sub/idx) format, could you add the 'character spacing' option,
to set in pixels the space/distance between characters?
Preferably, under the "Line height" option.
It is a really needed option when setting the look of the subtitles, especially when viewing them on TV.
http://oi63.tinypic.com/19ovsw.jpg
Nikse555
20th December 2018, 17:26
@orion44: The default libraries with c#/winforms don't have any 'character spacing' options as far as I know... anybody got any ideas?
Also, SE 3.5.8 is out - Now again includes Tesseract 3.02 per default and an in-program option to download the new Tesseract 4 (Tesseract 4 is pretty slow and do not have italic detection - but it has more languages + works better for small difficult fonts).
Also the translate works somewhat again...
Link to release page on github: https://github.com/SubtitleEdit/subtitleedit/releases
FLX90
7th January 2019, 15:43
Thank you for the great program. I used it for many BluRays and it worked like a charm.
Now I have a strange problem with the german Oldboy (2003) BluRay
I extracted idx&sub and the srt output would be:
1
00:00:50,800 --> 00:00:51,005
Was soll das?
2
00:00:51,009 --> 00:00:51,214
Was soll das?
3
00:00:51,217 --> 00:00:52,378
Was soll das?
4
00:00:52,802 --> 00:00:53,007
Hast du nicht verstanden?
Ich will nur mit dir reden.
5
00:00:53,011 --> 00:00:53,216
Hast du nicht verstanden?
Ich will nur mit dir reden.
6
00:00:53,219 --> 00:00:53,424
Hast du nicht verstanden?
Ich will nur mit dir reden.
7
00:00:53,428 --> 00:00:53,633
Hast du nicht verstanden?
Ich will nur mit dir reden.
8
00:00:53,636 --> 00:00:53,841
Hast du nicht verstanden?
Ich will nur mit dir reden.
9
00:00:53,845 --> 00:00:54,050
Hast du nicht verstanden?
Ich will nur mit dir reden.
10
00:00:54,054 --> 00:00:54,259
Hast du nicht verstanden?
Ich will nur mit dir reden.
11
00:00:54,262 --> 00:00:54,467
Hast du nicht verstanden?
Ich will nur mit dir reden.
12
00:00:54,471 --> 00:00:54,676
Hast du nicht verstanden?
Ich will nur mit dir reden.
13
00:00:54,679 --> 00:00:54,884
Hast du nicht verstanden?
Ich will nur mit dir reden.
14
00:00:54,888 --> 00:00:55,093
Hast du nicht verstanden?
Ich will nur mit dir reden.
15
00:00:55,096 --> 00:00:55,301
Hast du nicht verstanden?
Ich will nur mit dir reden.
16
00:00:55,305 --> 00:00:55,510
Hast du nicht verstanden?
Ich will nur mit dir reden.
17
00:00:55,513 --> 00:00:55,718
Hast du nicht verstanden?
Ich will nur mit dir reden.
18
00:00:55,722 --> 00:00:55,927
Hast du nicht verstanden?
Ich will nur mit dir reden.
19
00:00:55,930 --> 00:00:56,135
Hast du nicht verstanden?
Ich will nur mit dir reden.
20
00:00:56,139 --> 00:00:56,344
Hast du nicht verstanden?
Ich will nur mit dir reden.
21
00:00:56,347 --> 00:00:56,552
Hast du nicht verstanden?
Ich will nur mit dir reden.
22
00:00:56,556 --> 00:00:57,671
Hast du nicht verstanden?
Ich will nur mit dir reden.
Is this a bug on the BluRay?
I think it would be the best to do something like this:
1
00:00:50,800 --> 00:00:52,378
Was soll das?
2
00:00:52,802 --> 00:00:57,671
Hast du nicht verstanden?
Ich will nur mit dir reden.
But the output has 38564 lines.
Would be kind of sisyphean challenge.
I had to write a script.
Is there another workaround?
sneaker_ger
7th January 2019, 15:51
Tools->Merge lines with same text
FLX90
7th January 2019, 21:44
Oh, don’t saw that.
Thank you.
Awesome tool.
FLX90
7th January 2019, 22:12
Currently I'm using Tesseract 4.00 with engine mode Tesseract + LSTM.
Unfortunately I have to do massive post work.
'I' is sometimes a 'L' ,'l' or '!'.
Music symbol isn't recognized right (it's a P or D).
And I don't know what I have overlooked.
Is there anything I could improve in my settings?
EDIT:
I'm using binary image compare now and creating new database for every BluRay.
Works.
johner23
14th January 2019, 02:04
Hello, dear all.
I got the newest version from Subtitle Edit and try to convert a subtitle format to other.
The original subtitle is *.vtt fomat. The output is to be *.srt format.
But after the convertion the subtitle in srt show some strange errors.
Eg:
00:01:04,562 --> 00:01:08,908
{\an3}>> O que mais lhe
interessa na história?
The symbols {\an3} and >> are not correct.
Do you know how to fix that?
I can make manual correction after the convertion to srt, but I would prefer that the program could do it at the first attempt, without any "errors" or strange "symbols" after the task.
Thanks for your time.
Best regards.
Hendi
27th January 2019, 19:12
Hi
If I convert from sub to srt on batch-mode (/convert) the resulting srt file contains all timeframes but only the first timeframe contains a text. If I do this with the GUI, the whole text is converted and all timeframes contains text.
What do I wrong?
Thanks
Hendi
Nikse555
31st January 2019, 07:57
@Hendi: That's a bug, sorry - should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.8/SubtitleEditBeta.zip
Next final version should be out soon...
asarian
3rd February 2019, 01:55
Subtitle Edit 3.5 is now out
(...)
but SE can also [B]import and ocr vobsub and blu-ray image based subtitles (even from matroska/mp4 files), and DVB sub from .ts files
Subtitle Edit (the latest) actually immediately crashes for me when I import a .sup file ("Object reference not set to an instance of an object." and many other fatal errors).
Btw, why OCR-ing?! That's so 1985! :)
Nikse555
4th February 2019, 13:34
Subtitle Edit (the latest) actually immediately crashes for me when I import a .sup file ("Object reference not set to an instance of an object." and many other fatal errors).
Could you try latest beta?
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.8/SubtitleEditBeta.zip
SE 3.5.9 should be out soon...
dngnt
10th February 2019, 08:20
SE 3.5.9 should be out soon...
I'm using SE to convert from bitmap Chinese/Japanese/Korean subs (sup, idx/sub, DVDsub).
There is a possibility to save one sub picture at a time (with right click) while importing , but could you add a function of saving all the sub pictures during OCR?
For example, just a new option in the "export" (during OCR): export all the pictures as *.png? (just the images, no xml because it would be more complicated with "dirty" idx/subs (i.e with errors).)
It's because after finishing the OCR and saving the OCRed srt, I have to correct the (many) mistakes "offline" , but there is no way to compare the OCRed text to the original pictures.
Thanks for your great work!
Nikse555
10th February 2019, 13:40
@dngnt: In the OCR window you can right click in the list view... and choose "Save all images with HTML index..." or "Export -> BDN xml/png"
Nikse555
10th February 2019, 15:37
And SE 3.5.9 is out: https://github.com/SubtitleEdit/subtitleedit/releases
Some of the changes (mostly related to OCR) listed below:
* NEW:
* Bookmarks - thx OmrSi/marb99
* Image export - option to have single lines top justified - thx joedmartin
* IMPROVED:
* Improve Binary OCR of comma / apostrophe - thx Tuukka
* Improve quote/italic detection in binary OCR - thx Miggu
* Add context menu to OCR spell check
* Binary OCR auto detect best DB - thx Mr. Rage
* FIXED:
* Fix missing/bad html tags after "Auto br" - thx iromafia111
* Fix crash in OCR window when closing - thx spetragl
* Fix OCR in batch convert - thx danstraughn
* Fix crash parsing empty word in OCR via Tesseract - thx Barry
varekai
10th February 2019, 18:58
@ Nikse555
Thanks for the update, much appreciated!
Looking in the portable folders I see Tesseract302.
Is that version also included in SubtitleEdit-3.5.9-Setup.exe?
Kind regards
Nikse555
10th February 2019, 19:41
@varekai: Yes, Tesseract302 is included in the installer version as well.
Tesseract 4 is available as an in-program download, but note that T4 does not support italic detection and is a lot slower, but T4 is available in more languages and may work better for small unclear fonts.
varekai
10th February 2019, 20:57
@ Nikse555
Great! I use italics and only swedish and english languages so Tesseract302 works perfectly for my needs! Thanks!
dngnt
11th February 2019, 22:16
@dngnt: In the OCR window you can right click in the list view... and choose "Save all images with HTML index..." or "Export -> BDN xml/png"
Thanks for your indications!
Subtitle Edit is the best for its import/export versatility for srt, and with these 2 export abilities for sup,idx/sub and DVDsub, it's truly the Swiss knife for subtitling!
johner23
25th February 2019, 23:34
Hi, dear all.
@Nikse555
Hi, Nikse.
When I tested some files *.vtt I still get some errors. Different ones, I mean. Please, can you check it?
---> https://github.com/SubtitleEdit/subtitleedit/issues/3290
---> https://github.com/SubtitleEdit/subtitleedit/files/2873430/VTT.-.ERRORS.-.SAMPLES.zip
I can fix it manually using text editor, but I think it will be better if your program could do it automatically for us.
Thanks for your time.
Best regards.
Nikse555
28th February 2019, 22:04
@johner23: thx for the files - latest beta handles this better: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.9/SubtitleEditBeta.zip
SE adds/keeps the alignment though, so that why you'll see lines starting with e.g. "[\an8}". If you want to remove them, just select all lines in the list view (ctrl+a), right click, choose "Remove formatting -> Remove all formattings / Remove alignment".
locotus
1st March 2019, 00:23
Is there any way of sort lines by beginning capital or smal letters or
by lines beginning with small letters fallowing lines ending
with comma, space or other punctuation mark except period?
Thanks.
sneaker_ger
2nd March 2019, 11:36
Batch converting vobsub in mkv via cli /convert (lots of "Tesseract returned with code 1" messages, timestamps are there but lines are empty) as well as GUI Tools->Batch (lots of "Tesseract returned with code 1" messages, almost empty files) convert isn't working for me. I would also be nice to have a stream selector for mkv/mp4 input on the cli.
Nikse555
2nd March 2019, 16:38
@locotus: sorry, no - what exactly do you want to do?
@sneaker_ger: that should be fixed in last beta I hope - could you verify?
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.9/SubtitleEditBeta.zip (track-number parameter also added to cmd line)
locotus
2nd March 2019, 17:15
@locotus: sorry, no - what exactly do you want to do?
My main purpuse is to join dialogs lines that are
divided but with duration time and number of characters
below maximun range for both.
Actually I'm trying to do that sorting first for duration,
looking for very short lines with characters number greater
than 15. But can't do that with lines with longer time and
characters.
Thanks anyhow.
sneaker_ger
2nd March 2019, 17:28
Thx. I think there's a small typo: /? says "/trac-number:<track number>" (missing "k")
It still showed "Tesseract returned with code 1" over CLI but I figured out why: in the GUI OCR method was set to "Binary image compare". I changed it to "Tesseract 3.02" and the messages disappeared, the text appeared. I didn't expect the GUI setting to influence the CLI. Also it's not obvious why I would get Tesseract error messages if Tesseract isn't selected.
If you drag&drop mkv files into the GUI batch converter the GUI hangs for a long time, btw.
Nikse555
2nd March 2019, 20:10
@locotus: Could you use Tools -> Merge short lines?
@sneaker_ger: thx for testing :)
Beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.9/SubtitleEditBeta.zip
(I've not looked at freezing GUI in batch convert)
sneaker_ger
2nd March 2019, 20:21
Thx. Any changes besides the track/trac typo?
locotus
2nd March 2019, 20:49
@locotus: Could you use Tools -> Merge short lines?
The workflow I normally use is after running that step and after
lines unbreaker to make sure that when sorted for time duration the lines can be easy to visualize and fix after.
Rousseau
4th April 2019, 13:28
Is there any way to set the OCR language/dictionary during batch processing? It appears to use the last used settings instead of detecting language like it does in other program features.
18fps
16th April 2019, 09:36
Thank you for your great program! It has become my main tool for ripping subtitles.
I'd like to know if it's possible to (somewhat) keep the color information when ripping subtitles from blurays or dvds. Because I have many french blurays, and they don't use caps for HoH descriptions of noise or music, they use colors. Red subtitles are the ones that should be deleted for normal subtitles, yellow subtitles are off screen voices and sometimes should be converted to italics,and so on. So i'd like to be able to have a mark that identifies those colors for later processing. Maybe something like *r* for red subtitles or something like that...
nekrovski
5th June 2019, 07:34
Hi
I cannot seem to find a way to remove </span> tags with SubtitleEdit
https://i.imgur.com/kzPUIzp.jpg
I tried right clicking the line and choosing normal (remove formatting) and doesnt seem to work.
sneaker_ger
5th June 2019, 15:57
If there's no simpler option try regex via Edit->Multiple replace:
Replace what: <span.*?>(.*?)</span>
Replace with: $1
Select "Regular expression".
nekrovski
5th June 2019, 16:26
If there's no simpler option try regex via Edit->Multiple replace:
Replace what: <span.*?>(.*?)</span>
Replace with: $1
Select "Regular expression".
Thanks, that worked.
I've looked several times at all the options under "Tools" but cannot find a simpler method. It's weird, for a program like this to have every other little thing in tools, but not this. That's why I think I'm missing it somewhere.
kenzofabio
23rd June 2019, 15:34
hi everyone can u help me,i just translate with subtitle edit beta zip.but the result is some alphabet not a words,is there something wrong with sub edit.thx i appreciate for some information
mikmik68
8th July 2019, 07:49
Hi everyone
i have a problem with subtitle edit , when i translate from english to danish , there is a bunch of numbers and letters in a long string , before the correct translation ...
i have tried to delete everything including settings.
and reinstall the latest version.
nothing seems to work.
same result everytime, with the long string of numbers and letters in a long unbroken line instead of only the translation.
hope someone has a solution.
in advance , thankyou for you help :thanks:
Regards Kim
Nikse555
16th August 2019, 13:37
mikmik68: The issue with random numbers/letters in GT (Google Translate) should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.9/SubtitleEditBeta.zip
Also, SE 3.5.10 should be out soon, so please give the above beta a test run :)
sneaker_ger
18th August 2019, 18:01
Could you make ASS/SSA more compatible with existing software/"standards"?
1. some software can't read ASS/SSA when there isn't a Newline at the very end of the file (like Aegisub writes)
2. you use the field "Actor" instead of "Name" for ASS. The "spec" and Aegisub use "Name". I think some user requested "Actor" but using "Name" can serve the same purpose just more compatible.
Nikse555
19th August 2019, 16:44
@sneaker_ger: thx for the info - I've tried to add 1+2 in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.10/SubtitleEditBeta.zip
Let me know how it works.
Also, SE 3.5.10 is out :)
sneaker_ger
19th August 2019, 17:50
Beta is looking great. Thank you for your continued work!
sneaker_ger
5th September 2019, 22:33
(Latest Beta regex:) Why is ".*" replaced by "aa", not "a"?
https://i.imgur.com/Fn5V2or.png
Nikse555
8th September 2019, 08:17
It seems that ".*" also matches an empty string. So first "line1" will be replaced by "a" and then "" will also be replaced by "a" too giving "aa".
See https://stackoverflow.com/questions/7326679/net-regex-matching-matches-empty-strings
So try ".+" instead to avoid empty strings... I don't know why it works like that ;)
sneaker_ger
8th September 2019, 09:18
Ah, I see. It's different from other software I use (e.g. Sublime Text). Thank you.
Matt Kirby
17th September 2019, 19:17
Hi,
I have a problem with OCR and TS-Files.
https://www.dropbox.com/sh/vxsz74el242qw5t/AABYB-56jzODmDPgZlVrxHIaa?dl=0
Here is a part with exact one subtitle-item. When I load it in SE the length is 3.5 seconds.
That's too short. When I convert the part.ts with mkvtoolnix to part.mkv
and I load it to SE the length is correct (about 10.xxx seconds).
Please could you change SE so that the correct length is given with TS-files without conversion?
shag00
18th September 2019, 11:52
I know this request has been going on since 2011 but it really is time this software was able to be easily used by linux users. If an appimage, containing VLC (https://www.videolan.org/legal.html) could be made available it would make life very much easier for linux users. The link on your download page for linux users is 8 years old and out of date. Although someone, somewhere may have cracked it, as far as I am aware you cannot get a windows quality experience with SE using either wine or mono. Either the OCR or the video always has some problem, and stability is also an issue. Of the mastering that I do, sound is now usable decoding, editing and reassembling in linux, video is still outstanding as is subtitles. There is really no native linux alternatives, what is available seems to be no longer supported and certainly is not as fully featured as SE.
Ghitulescu
18th September 2019, 13:58
With so many text-processors, sinking this software into an infinite sea of requests is a bit counterproductive. I mean, when two requirements of treating the semicolon or the hyphen exists, why these people do not use their specially-tuned text-processors, adapted for their language, and do all the text OCR/proc and let SE do only the actual subtitle editing.
I've seen this syndrome eg with satellite receivers whose owners asked the manufacturers or the hackers to fine-tune the AVI-playing section.
Is there any improvement in the bitmap handling of SE in the last 5 years?
shag00
18th September 2019, 23:20
With so many text-processors, sinking this software into an infinite sea of requests is a bit counterproductive.
Is this in response to the above post or just a random thought on people asking for product enhancements?
Ghitulescu
19th September 2019, 08:43
Is this in response to the above post or just a random thought on people asking for product enhancements?
This is a gentle reminding of a feature I've asked some 5 years ago, thinking that maybe the author got caught in a repetitive-loop. At that time it was the only software that could deal with images in BD/M2TS format.
OCR is performed OUTSIDE of this software, and if one half of the users require a -dash before the first line of dialogue and the other half require exactly the opposite, and this goes on and on for the aforementioned period of time, well, ... :)
Nikse555
20th September 2019, 21:08
@Matt: Atm, if the duration is longer than "max duration" in settings, then SE just set the duration to 3,5 secs - just changed the code to use "max duration" instead of 3,5 seconds - but perhaps it would be better to just let the user fix that afterwards?
You can set the "max duration" to something large like 20 seconds to allow long duration.
@Ghitulescu: Most changes to SE are from people asking for product enhancements, but I do really try to be careful about what features I add - and I don't want to add everything to SE (like advanced ASS features as we already have Aegisub for that or features that will have very few users). SE focuses on creating/adjusting/fixing/reviewing text subtitles and ocr'ing/converting + translating - the last one is ever changing :(. New subtitle formats are also very important (like new stuff from Amazon, Netflix or w3c).
In the last five years many improvements has been made in reading bitmap formats from mp4/ts/m2ts/bdsup and converting between/to bitmap based formats - read more here: https://raw.githubusercontent.com/SubtitleEdit/subtitleedit/master/Changelog.txt
What was your original request by the way?
Matt Kirby
21st September 2019, 11:22
@Matt:
You can set the "max duration" to something large like 20 seconds
Thank you! That works...
Ghitulescu
23rd September 2019, 10:51
What was your original request by the way?
I kindly asked for a possibility to change the colours (maps) of the SUP, sort of what DVDSubEdit did for DVD. Essentially, I need to get rif of the stupid black box surrounding the text, to make it transparent.
Nikse555
28th September 2019, 07:33
...I need to get rif of the stupid black box surrounding the text, to make it transparent.
Could you post/link/email a sample subtitle?
Ghitulescu
2nd October 2019, 21:35
Could you post/link/email a sample subtitle?
I don't know how to cut full streams comprising subtitles (and to keep them in), therefore I needed 3 days to delete 2000 images one-by-one in order to fit the 200kB limit of doom9:
nekrovski
6th October 2019, 10:36
Is there a way to fix capitalizion on next line when there wasn't a fullstop/question mark/exclamation etc on the previous line?
https://i.imgur.com/Tc3Ofp1.jpg
Nikse555
6th October 2019, 14:19
@nekrovski: Try Tools -> Change casing... Normal casing
You can do a compare afterwards to see the changes.
Nikse555
8th October 2019, 06:53
@Ghitulescu: Perhaps you can use some file host for the sample file or email it to me?
In Bluray sup files each subtitle have it's own palette (up to 255 colors), and vobsub normally has a global 4 color palette.
Ghitulescu
8th October 2019, 14:58
In Bluray sup files each subtitle have it's own palette (up to 255 colors), and vobsub normally has a global 4 color palette.
I know this :(
nevertheless it's easy to find out the background color index entry.
nekrovski
9th October 2019, 09:37
@nekrovski: Try Tools -> Change casing... Normal casing
You can do a compare afterwards to see the changes.
Thank you, this worked.
Verminaard
10th October 2019, 14:22
I have a question: When I use "remove text for hearing impaired" it removes the unnecessary SDH tags fine but it adds speaker dashes to some lines. What could be the reason to this? Version 3.5.10.
Nikse555
11th October 2019, 06:04
@Ghitulescu: Sorry, the attached subtitle is pretty impossible to do anything about... sometimes the text is completely transparent and sometimes the text is just as transparent as the box.
@Verminaard: Could you give some examples? Speaker dashes should be added sometimes...
nekrovski
11th October 2019, 09:22
I opened the non hearing impaired subtitles from the new Breaking Bad movie, and it didn't find any common errors.
This is the first time I see an original subtitle in which Subtitle Edit doesn't find any common errors.
I was appalled.
Did they, by any chance, used Subtitle Edit for the subtitles? :D
Nikse555
11th October 2019, 09:38
@nekrovski: that may be possible: https://partnerhelp.netflixstudios.com/hc/en-us/articles/115000258712-Subtitle-Edit
Ghitulescu
11th October 2019, 17:45
@Ghitulescu: Sorry, the attached subtitle is pretty impossible to do anything about... sometimes the text is completely transparent and sometimes the text is just as transparent as the box.
It's one of those dynamic pallettes, I know, but it's very simple to identify and then change only the background.
Verminaard
11th October 2019, 22:23
@nikse55
I will send you a link for an example file later. But logically, a function called "Remove" should not be adding something now should it :) Fix subtitles on the other hand has an option to add dashes. But in my case they're unchecked.
Thanks for responding.
von Suppé
14th October 2019, 22:36
I have the same problem as Verminaard.
But, after some testing, I think I found something. The tool may have an issue when text for hearing impaired are BOTH italic AND in double lines. I created a little .srt file (attachment) that should explain.
dngnt
17th October 2019, 20:57
@dngnt: In the OCR window you can right click in the list view... and choose "Save all images with HTML index..." or "Export -> BDN xml/png"
I've been using this feature quite a lot. I would suggest that when using
"Export -> BDN xml/png" --> Export all
instead of presenting Desktop as default export directory, to use the current location of the idx/sub or sup file ?
That would be a big time saver.
Thanks for your great job!
Nikse555
18th October 2019, 05:04
I've been using this feature quite a lot. I would suggest that when using
"Export -> BDN xml/png" --> Export all
instead of presenting Desktop as default export directory, to use the current location of the idx/sub or sup file ?
That would be a big time saver.
Thanks for your great job!
@dngnt: thx for the info :)
I've tried to fix it here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.10/SubtitleEditBeta.zip
Does that work for you?
Nikse555
23rd October 2019, 21:13
I have the same problem as Verminaard.
But, after some testing, I think I found something. The tool may have an issue when text for hearing impaired are BOTH italic AND in double lines. I created a little .srt file (attachment) that should explain.
Thx for the info - latest beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.10/SubtitleEditBeta.zip
von Suppé
24th October 2019, 09:47
Hi Nikse,
When I load the test srt (which I uploaded) in your latest beta, the issue of adding double dashes is fixed, indeed. But it still adds a dash to the first line?
Zetti
27th October 2019, 11:41
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.11
Nikse555
3rd November 2019, 15:05
@von Suppé: SE prefers dialog dashes for both lines... but please do post examples if you find bugs :)
@Zetti: you're welcome, seems to be a nice and stable release!
Also, if you have transport streams (.ts) files with teletext, please do test latest beta version (use File -> Open.. and choose .ts file): https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.11/SubtitleEditBeta.zip
dngnt
3rd November 2019, 19:03
@dngnt: thx for the info :)
I've tried to fix it here: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.10/SubtitleEditBeta.zip
Does that work for you?
I've tried your latest Beta and your 3.5.11 version with the export --> png/xml .
it now offers "Desktop" for default for the output.
Still, I would prefer the current location of the .sup or idx/sub file ( which is on a separate hard disk, and a specific directory) for the default.
When I import from...and then do the OCR and then save... the current sub file location is by default.
With the export feature, one should expect the same behaviour .
Thanks a lot for your versatile tool!
von Suppé
17th November 2019, 15:35
Hi Nikse,
I am figuring out how to to deal with some issues I run into when using another subtitle tool, regarding SUP --> XML/PNG conversion and vice versa. In doing so, but also for just wanting to know this now, I ask myself, how are BD SUP files build up?
Since SubtitleEdit can export to BD SUP, I wonder if you can make some things clear.
As for graphics, of course I can export SUP to XML/PNG and take a look at the .png images. Doing this with an image-editor, it tells me that it's "2D" (only 1 layer) and of course it carries opacity-data.
Again, this is after SUP --> PNG/XML conversion. Are the pictures in SUP indeed stored as "2D" PNG (or uncompressed BMP32) or is there more to them?
Also (this would be ages ago) I read something about the picture's I/O timings being defined differently, using framenumbers & framerate instead of time-codes such as used in SRT. Never giving it more thought at the time, but now it's more topical for me.
I hope you can shed more light - for a newbie, as this is the first time I go more deeper into SUP.
Nikse555
20th November 2019, 12:26
@von Suppé: Some nice info about bdsup is available here: http://blog.thescorpius.com/index.php/2017/07/15/presentation-graphic-stream-sup-files-bluray-subtitle-format/
bdsup does not have frame number (it does have a "frame rate" but that is mostly not used), so it's time based like SRT.
Each image has a palette with 255 colors and each entry also has an "alpha" (lumianance) - much like an 8-bit png.
A bdsup entry can have multiple images (but mostly don't)
von Suppé
20th November 2019, 12:50
Thanks for the info, Nikse. I'll go read.
darksen
2nd December 2019, 00:09
I've been using your program for a very long time and it is great. I always wanted to make a request, but I never registered here until some time ago.
I make a big use of the replace list with a lot of custom regex and other stuff and as such I have set it the best I could so I don't need/have to review every little change it does. My request is to split the message: "Fix common OCR errors (using OCR replace list)" into two messages, one pointing that the change is made using the default replace list and other pointing that the change comes from the user replace list. That way I can disable the changes I have to review and apply the changes I know I don't have to check (the ones coming from the user replace list) so in the second pass I go one by one checking them, because a lot of times I got more than 200 hundred entries in there and going one by one checking what has changed to disable the ones that shouldn't be applied is very time consuming, sometimes I don't have to disable anything because all comes from my user list so I just waste time checking each entry.
Please, if you can consider adding this it would mean a lot.
Thank you.
nekrovski
5th December 2019, 10:18
Hi,
I was working on a subtitle. Then I used Handbrake to burn in the subtitle in the video.
When I played the video, this showed up
https://i.imgur.com/I9BNR0H.jpg
When the video is played without burned in subtitles, but instead with subtitles from a separate .srt file from the video, it shows fine.
I opened the subtitles with Notepad and I noticed that this empty space highlighted here
https://i.imgur.com/fXB3ydT.jpg
confuses Handbrake.
Is there a way for Subtitle Edit to detect errors like this with the Fix common errors option?
I tried but couldn't find.
Nikse555
19th January 2020, 15:52
SE 3.5.12 is out :)
Can now read teletext from .ts / .m2ts files - includes any colors + top alignment (.ts reading should also be faster)
Tesseract updated from 4.1.0 to 4.1.1 + more pre-processing image settings
@darksen: Did you get a lot of false corrections? Perhaps that could be improved if you posted some.
@nekrovski: It's not really errors with whitespaces... you could re-save with SE via command line perhaps?
EDIT and off topic: Is it possible to change the thread name? I've tried to change the title in the first post, but the thread title still says "Subtitle Edit 3.5.10"...
mkver
4th February 2020, 15:23
Now that I wanted to test the new teletext feature I directly stumbled upon the issue behind my ticket 1897 (https://github.com/SubtitleEdit/subtitleedit/issues/1897) again: It should be possible to treat several transport stream files as one big file (in case a DVR only supports FAT32 and therefore splits files into pieces smaller than 4GB each).
arslan
5th March 2020, 05:02
Hi Nikse,
would it be possible to implement the new Serbian dictionary and spellchecker, which is significantly more comprehensive than the current one?
(The new one is located here:
https://github.com/msmiljan/korektor
and here:
https://extensions.libreoffice.org/extensions/serbian-spellcheck-and-hyphenation)
Thank you very much!
GCRaistlin
17th March 2020, 17:04
How can I use portable MPC-HC with Subtitle Edit? MPC-HC option is greyed out in Settings.
Feature requests:
[Options - Settings... - Word lists] Double click on a pair in OCR fix list fills the fields beside 'Add pair' button with the corresponding values.
[Options - Settings... - Tools - Fix common OCR errors - also use hard-coded rules] Make using hard-coded rules customizable. For example, replacing 'l' between uppercase letters with 'I' is surely needed while converting the first letter of the paragraph to uppercase may be completely unwanted.
[Import/OCR Blu-ray (.sup) subtitle file...] Add the ability to disable spell checking while still using OCR fix list for selected language. This makes sense because some errors like "l instead of I after the dot" aren't being fixed by OCR fix list and hence force spell checking dialog to appear. But they may be fixed by applying a regexp in an external editor (for the mentioned error it would be "(?<!\w)l(?!\w)"). Applying regexps before spell checking saves a lot of time but to use regexps currently we need to OCR without error fixing and then call Fix common errors tool with only 'Fix common OCR errors (using OCR replace list)' option checked.
Add the ability to use regexps to fix common OCR errors. It would be great to create a predefined set of regexps like the one above. I'm ready to share my own.
Look for Settings.xml in the current (working) directory (i. e. directory that was the current when SE was laucnhed) instead of the directory where SubtitleEdit.exe is located. It would allow to have different settings for different cases or users.
Bugs (Import/OCR Blu-ray subtitle, OCR method: Binary image compare, Image database: Latin):
Non-Italic dashes that are followed by Italic text are erroneously recognized as Italic (example (https://i111.fastpic.ru/big/2020/0317/1e/2bdc9546c5f53f9260e09643bcd76c1e.png)). Also another bug with this example subpic: if Dictionary field is empty then the space after "Audience" is lost; if Dictionary is set to English then the space is preserved.
"t ]" in Italic is recognized as "t]" with default "8 pixels is space" (example (https://i111.fastpic.ru/big/2020/0317/b1/4aa295f7e8614bf5a559e097361f38b1.png)).
'9' is recognized as '0' (example (https://i111.fastpic.ru/big/2020/0318/ce/a06f2c687da4c65aadd7f52b14a772ce.png)).
Jumping to a subpic by typing its number in 'Subtitle text' area isn't working for #555: typing '5' repeatedly moves the cursor from #50 to #51, then to #52... #59, then to #500 and so on.
'Fix OCR errors' checkbox state isn't being saved.
OCR fix list (English):
Why default OCR fix list includes "backseat -> back seat"? Is "backseat" really incorrect?
Why default OCR fix list includes "lt -> it"? I believe it should be "lt -> It". The same thing with various lf-started pairs: one part of them has "If" as a result (which is correct), another part has "if" as a result (which is not).
Lucius Snow
18th March 2020, 12:53
Hi all,
Since the update to 3.5.13, when I reload existing subtitles from a recent SRT, it doesn't load anymore the video which goes with it.
Can you please tell me how to restore this?
Thank you.
Nikse555
18th March 2020, 15:15
@arslan: thx, will be included in next update.
@GCRaistlin: A lot of input...
You probably need a 64-bit version of MPC-HC... but I would recommend that you try "mpv" as video player, as that seems to be the best option atm.
You can use "Ctrl+G" for go to sub#
'9' is recognized as '0'... Double click on the line in the list view, and use "Add better match".
5) Look for Settings.xml... that's what profiles were made for.
Regular expressions are already supported - see the english ocr fix replace list.
"backseat -> back seat"... yes, that seems like a bug, thx.
@Lucius Snow: SE 3.5.14 is out - and try "mpv" as video player (Options - Settings - Video player - Download mpv lib)
GCRaistlin
19th March 2020, 09:59
You probably need a 64-bit version of MPC-HC
I have both 32-bit and 64-bit portable versions installed. How SE recognizes that MPC-HC is (not) installed?
'9' is recognized as '0'... Double click on the line in the list view, and use "Add better match".
That's what I've already done but isn't it an everyone's issue?
that's what profiles were made for.
Profile on General tab? I don't see how to add a new one there. Also, changes seem to be written to the current profile without asking. For example I've changed 'Single line max. length' and the new value has been saved to 'Default' profile immediately.
Regular expressions are already supported - see the english ocr fix replace list.
I don't see any regexps there (Words lists - OCR fix lists). Also, I mean support for manually applying regexps, not without asking.
Remove text for hearing impaired issues:
It doesn't remove commas: "I am, uh, late" -> "I am, late". It is clear that it may be necessary commas but to leave them all - without the possibility to edit text right there - isn't a good decision, too.
It's hard to identify a caption that needs to be edited after using this tool: only caption numbers are showed but the numeration will be changed due to the deletion of captions. So the only way is to write down a portion of the target caption text and then search it. Displaying of time codes (or/and log of applied fixes available after quitting the tool) would be great.
Now in List view Start time and Duration are displayed and available for editing. In some cases, editing of End time is preferable. Can you please add such a possibility?
The biggest trouble for me is though non-customizable hard-coded rules applying. To prevent making first letters uppercase I'm forced to perform a global replace with a regexp before Fix common errors and another global replace after.
Lucius Snow
19th March 2020, 18:48
@Lucius Snow: SE 3.5.14 is out - and try "mpv" as video player (Options - Settings - Video player - Download mpv lib)
Thank you.
I already use mpv but I installed it manually because it doesn't work from Subtitle Edit (error from a server with no SSL/TLS).
Same with 3.5.14.
Nikse555
19th March 2020, 21:49
@GCRaistlin: Do have have some examples where first letters are converted wrongly to uppercase via hard coded rules?
About the profile on General tab... click the "..." button the make new or delete profiles.
The double-click on word lists feature are available in latest beta.
About MPC-HC, you can use the "MpcHcLocation" in settings.xml or just a subfolder called "MPC-HC" in the SE folder, using the installer should also work (MPC-HC does not really have an API - SE actually just steals the video from the MPC-HC UI) - but just try mpv :)
@Lucius Snow: Hm, that works fine here... how can I re-create your issue with video not opening? Can you give more information? Is it all .srt files or only some. Video type? Do you have the video window open?
Nikse555
19th March 2020, 22:34
@GCRaistlin: Editing the text in "Remove text for HI could be possible - how does this work: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
Lucius Snow
19th March 2020, 22:52
@Lucius Snow: Hm, that works fine here... how can I re-create your issue with video not opening? Can you give more information? Is it all .srt files or only some. Video type? Do you have the video window open?
Anyway, mpv seems properly installed because I use it to play videos. I manually installed it. The video does play.
My problem is just when I open an existing SRT from File / Reopen menu. Before the update, it opened both the SRT and the associated video file. Now, it doesn't open the video. I have to do it again each time I open the SRT.
It concerns any codec / container.
Nikse555
20th March 2020, 07:08
@Lucius Snow: Do you still have problems in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip ?
If yes, what are the steps for re-creating this in detail (drag-n-drop or file open or shortcuts)?
GCRaistlin
20th March 2020, 11:54
some examples where first letters are converted wrongly to uppercase via hard coded rules
92
00:07:54,641 --> 00:07:56,559
- l would take the idea to its extreme -
- [ Whispering, lndistinct ]
93
00:07:56,643 --> 00:08:00,188
and draw parallels
between reproduction in art. . .
577
00:38:09,245 --> 00:38:13,875
- Well, he said that, uh -
- lt is actually as beautiful as the original.
578
00:38:14,000 --> 00:38:17,295
- that they thought it was an original
for many, many centuries -
- [ Man, ln ltalian ] When was it made?
1091
01:12:28,344 --> 01:12:32,014
- That impression is quite right, but. . .
- [ Crowd Cheering ]
1092
01:12:32,181 --> 01:12:33,808
how can l say. . .
See ## 93, 578, 1092. Anyway, "OCR error" means that something is erroneously recognized while the first letter may be in lower case in an original caption. It may be an error in a general sense, but if we just want to get the text that is identical to the graphical source we surely don't want such AI.
About the profile on General tab... click the "..." button the make new or delete profiles.
Oh I see, how could I just miss it. But searching for Settings.xml in the current directory first would be still useful for multi-user environment.
The double-click on word lists feature are available in latest beta.
Working, thanks. Can you please implement replacing of an existing word on 'Add pair' press (with confirmation)?
Editing the text in "Remove text for HI could be possible - how does this work
It does, thanks again.
Lucius Snow
20th March 2020, 13:08
@Lucius Snow: Do you still have problems in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip ?
If yes, what are the steps for re-creating this in detail (drag-n-drop or file open or shortcuts)?
Thank you but there's no change.
Nikse555
20th March 2020, 13:35
Thank you but there's no change.
But can you give steps to re-create this issue?
Lucius Snow
20th March 2020, 14:29
But can you give steps to re-create this issue?
That's what I described earlier:
My problem is just when I open an existing SRT from File / Reopen menu. Before the update, it opened both the SRT and the associated video file. Now, it doesn't open the video. I have to do it again each time I open the SRT.
Difficult to explain more :(
GCRaistlin
20th March 2020, 22:28
Nikse555
Now in List view Start time and Duration are displayed and available for editing. In some cases, editing of End time is preferable. Can you please add such a possibility?
In addition to my previous request I'm offering to add 'Pause before next' field. The result could look like this:
[x] Start time: ___ [ ] Duration: ___
[ ] End time: ___ [x] Pause before next: ___
The idea is that 0, 1 or 2 checkboxes can be set at the same time. Inactive fields (related to cleared checkboxes) are greyed out, their values get changed in accordance to the values in active fields.
This bug isn't reproducible with Latin.db in the latest beta but I decided to report it 'cause it is really strange. Try to OCR these SUP(BD) subtitles (https://mir.cr/0RLUHRRR) with 3.5.14 (clear all checkboxes but [x] Fix OCR errors). The problematic caption is #203. If you start from #100 (skip unrecognized characters twice) "Sunday" will be recognized as "sunday". If you start from #101 (close current OCR session and start a new one) it will be recognized as "Sunday".
Currently, the installation package contains files that may be changed by an user (Latin.db, en_US_user.xml and so on). Hence, they may be replaced with default ones by update. It would be better if all changes were made to the files that don't exist by default.
My current Latin.db (https://mir.cr/19P4ZIDN) - maybe you'll find my additions useful for all.
tormento
21st March 2020, 10:31
@Nikse555
I am doing some OCR on idx+sub files and every "I" that begins a sentence is converted to "L".
Can you fix that?
Here (https://www.upload.ee/files/11307246/zcd.7z.html) is a sample.
tormento
29th March 2020, 11:31
@Nikse555
I am finding some encoding giving problems to your editor.
Here (https://www.mediafire.com/file/kkr62qfk1mvpica/SupRip_samples.zip/file) you can find some. I hope the names are self explanatory enough.
The only ones I can open with no problems on accented vowels and symbols are UTF16-LE ones. With UTF8 it is a mess ;)
Nikse555
29th March 2020, 12:08
@tormento: I don't think those files are correctly encoded, sorry.
EDIT: The UTF-8 files have UTF-8 BOM (EF BB BF) but they are not using UTF-8 encoding, they are ANSI encoded! Yes, really a mess :)
tormento
30th March 2020, 12:28
@tormento: I don't think those files are correctly encoded, sorry. EDIT: The UTF-8 files have UTF-8 BOM (EF BB BF) but they are not using UTF-8 encoding, they are ANSI encoded! Yes, really a mess :)
They come from Sub Rip, latest version. I dunno if author is still active to report him this mess.
What about the other message about "L" OCR?
Nikse555
30th March 2020, 14:41
What about the other message about "L" OCR?
I've updated latest beta somewhat: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
Your subtitle runs very well through the OCR using "Binary image compare" with number-of-pixels-is-space=7 and max-error-pct=1.
I did not have any problems with "L".
What OCR method are you using and what lines are problematic?
tormento
31st March 2020, 12:57
Your subtitle runs very well through the OCR using "Binary image compare" with number-of-pixels-is-space=7 and max-error-pct=1. I did not have any problems with "L".
Clean installation, same settings of yours. I drop idx to SE, no OCR auto correction enabled.
Many many "L":
570
00:51:39,520 --> 00:51:42,478
L know the book is tough,
but l liked it.
571
00:51:42,600 --> 00:51:43,476
L know.
Tried with a fresh installation and "untrained" OCR database?
Nikse555
31st March 2020, 14:32
Clean installation, same settings of yours. I drop idx to SE, no OCR auto correction enabled.
Many many "L"...
I get "l" (lowercase L) instead of "I" (uppercase i)... because the two images are exactly alike. Enabling "Fix OCR errors" should fix those...
Result here, starting with lowercase "L":
l know the book is tough,
but l liked it.
tormento
31st March 2020, 15:20
I get "l" (lowercase L) instead of "I" (uppercase i)... because the two images are exactly alike. Enabling "Fix OCR errors" should fix those...
Result here, starting with lowercase "L":
l know the book is tough,
but l liked it.
I get capital L!
Boulder
2nd April 2020, 18:15
I get capital L!
I've added the OCR fix list pair (Options -> Settings -> Word lists) l --> I to fix this, if I remember correctly. You can also do that after the OCR run.
That subtitle example was very straightforward to OCR and the characters look like most DVDs, so you get a lot of good matches for future subs.
I uploaded my dictionary files and latin.db in case someone finds them useful (Nikse555 can freely use the content with SE if he wants to):
https://drive.google.com/open?id=1BoeF5_dwzIVpbxJYWS8tAwXRn-IWRgnz
https://drive.google.com/open?id=1Nz1JLOmhO8PIYajV5RCBOW9K93DZVxHb
tormento
3rd April 2020, 11:07
One of the best things of SubRip was the possibility to save different matrixes and automatically scan for the most effective one on OCR when loading the sub bitmap file.
That gives the possibility not to pollute the good trained ones with some unusual subtitle, plus the possibility to organize them effectively.
Moreover, Subtitle Edit could come with pretrained ones like SupRip did for the most commonly used fonts.
@nikse would you, please?
Boulder
3rd April 2020, 11:14
SE does have the ability to use different databases for OCR and the scanning could be useful, I agree.
tormento
3rd April 2020, 11:21
SE does have the ability to use different databases for OCR and the scanning could be useful, I agree.
Another missing thing is the possibility to expand selection, such as for % that sometimes goes wrong on OCR. A point and click expansion thing would be even better.
Boulder
3rd April 2020, 11:37
If the character is not recognized, it's possible to expand. If it's detected wrong, afterwards, still in the OCR dialog you can right click on the line with the issue and choose to inspect the matches. Then right click on the incorrect match and you get an option to select a better multi match. It's something I reported as an issue a long time ago so I just happen to know where it is. It's not easy to come by by accident, the same with all those special characters that you can add using the right click in the OCR dialog where it asks which character the image represents. I've used the software for years and found this one out last week :)
tormento
3rd April 2020, 12:05
If the character is not recognized, it's possible to expand.
How?
Manually changing the OCR recognition is something I already know.
Expanding the bitmap is new to me.
Boulder
3rd April 2020, 12:08
How?
Manually changing the OCR recognition is something I already know.
Expanding the bitmap is new to me.
There should be the button to expand selection if the process runs into a character it does not recognize. Off the top of my head, it's in the top area of the window. Using it is so automatic to me that I don't recall the exact place.
EDIT: the problem remains if the first part of % (a dot) is recognized, expansion only works forwards and not backwards. In these cases, I abort the OCR and check the matching for that line manually concerning the incorrect detection and restart the process from there.
tormento
3rd April 2020, 12:13
There should be the button to expand selection if the process runs into a character it does not recognize.
Where?
https://i.lensdump.com/i/jkrp7z.md.png (https://lensdump.com/i/jkrp7z)
Boulder
3rd April 2020, 12:15
It's not in that main dialog. It appears when you start the OCR process, in the bitmap/character matching phase if there is no match.
tormento
3rd April 2020, 12:17
It's not in that main dialog. It appears when you start the OCR process, in the bitmap/character matching phase if there is no match.
Yep, you are right.
Boulder
3rd April 2020, 12:19
And if you already haven't: set/download and set the correct dictionary and enable Fix OCR errors to make your life a bit easier.
tormento
3rd April 2020, 12:40
And if you already haven't: set/download and set the correct dictionary and enable Fix OCR errors to make your life a bit easier.
I tried to use it but it doesn't work really well for italian.
Boulder
3rd April 2020, 13:15
I tried to use it but it doesn't work really well for italian.
There's two dictionaries, are the equally bad? You can of course modify the dictionaries if there are clear errors. In my opinion, they are the key to getting the whole process fast and accurate but it takes a lot of time in the beginning.
tormento
4th April 2020, 11:14
Ok, I reset Latin.db and started to create a better OCR database, using italic and so.
I think the fact to have to reopen the same IFO for different subs different times is really annoying.
Am I doing something wrong?
jlw_4049
15th April 2020, 19:18
Anyway to minimize the program while it's OCR'ing?
Also is there anyway to make the program flash when it has a prompt on the task bar, so you know?
Also last question, when updating from older versions, where is the user dictionary/name dictionary saved? That way it's carried over to the newer version?
Also, is there anyway to minimize the OCR window? To let it do it's work in the background instead of being forced above all other applications?
Thanks! Program is awesome!
Nikse555
16th April 2020, 10:37
@tormento/Boulder: A few years back I actually tried to programmatically create a Binary-image-compare-db with all windows fonts in 5 different sizes... it's was slow and to my surprise not even very good.
I'm not really sure what the best solution is, but I'm pretty sure it's not one db.
Binary image compare and "backward expansion": If you have a difficult character (like "%" where first part is recognized as "o") - you can fix it in the "Inspect compare matches" (dbl click on line in list view), right click in "Inspect items" and choose "Add better multi match".
@tormento:
>I think the fact to have to reopen the same IFO for different subs different times is really annoying.
In the "Choose language" window you can do a "Save as..." for each language stream id.
@jlw_4049:
To minimize the program while it's OCR'ing: latest beta now has enabled the minimize icon.
Latest beta blinks in the taskbar when OCR has a prompt (up to 25 times, but only if OCR window is not focused).
Dictionaries are saved in... press Win+R, paste "%appdata%\Subtitle Edit\Dictionaries" and press enter. Most (*not all*) dictionaries have a "_user" file... e.g. the Binary-image-compare-db "latin.db" does not.
Please test latest beta :)
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
(can also use "Tesseract 5 alpha")
tormento
16th April 2020, 11:55
@tormento/Boulder: A few years back I actually tried to programmatically create a Binary-image-compare-db with all windows fonts in 5 different sizes... it's was slow and to my surprise not even very good.
Not necessary to build for every single font. Subtitles manly use arial/helvetica derivatives and for the "strange ones" we can build db on our own. That's why I suggested you to use more than one db, so we can add more characters without "polluting" our standard font sets. Please give a look to SupRip and SubRip. Boh uses more than one set of fonts and the last has the ability to scan thru them to recognize the most fitted.
Nikse555
16th April 2020, 13:54
SE already scans through dbs and chooses the best fit (I hope).
I guess that similar fonts should have their own db for each font size - to avoid "polluting".
What fonts should be used?
tormento
16th April 2020, 15:24
I hope
It's you the programmer! You should know! :D
What fonts should be used?
SupRip uses (you can find the file inside, did you look?):
arial-bold.font.txt
arial.font.txt
arial2.font.txt
arial3.font.txt
arial4.font.txt
calibri-variant.font.txt
geneva-bold.font.txt (I think it's Helvetica)
greek-arial.font.txt
narrow.txt (I think it's Arial Narrow)
SubRip uses cryptical file names but there are 106 font matrixes.
The least could be to start create multiple databases on our own, with your assurance that SE scans for them or it would be useless.
Nikse555
16th April 2020, 17:59
@tormento: SE does search for the best match of OCR DBs (only via first sub now, but that could easily be changed)
jlw_4049
16th April 2020, 18:01
@jlw_4049:
To minimize the program while it's OCR'ing: latest beta now has enabled the minimize icon.
Latest beta blinks in the taskbar when OCR has a prompt (up to 25 times, but only if OCR window is not focused).
Dictionaries are saved in... press Win+R, paste "%appdata%\Subtitle Edit\Dictionaries" and press enter. Most (*not all*) dictionaries have a "_user" file... e.g. the Binary-image-compare-db "latin.db" does not.
Please test latest beta :)
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
(can also use "Tesseract 5 alpha")
Thank you! I'll try it out and let you know! :)
Once minimized it cannot be re-opened until it's done. Which isn't a negative for me, but just letting you know!
Edit:
Is there anyway to do a batch OCR. I had 10 copies of the program last night and it had a memory leak and I had to hard power cycle it.
That way I can add multiple files at a time and they get done 1 by 1?
eddified
20th April 2020, 03:53
I tried this app out for the first time today. Seems pretty awesome. Thank you so much! My first try reading PGS, I went with default OCR settings, binary image compare (don't do that, results were terrible -- just about every "o" was interpreted as a G). Though almost all text was italic, maybe that was the problem with the bad "
The very first time, it asked me to confirm character by character. Ex: it showed me an "a" and said, "what is this?". I just clicked ok for awhile. It took me a few subtitles before I realized I was telling it every character was "" (the empty string) in text. Then I wondered if I had ruined some dictionary by having it save away those bad values. So I ended up starting again from scratch -- which requires a time-consuming re-load. By the 3rd try I realized Tesseract was much, much better than the default binary image compare. :)
Some questions:
1) Loading a Blu-ray rip file takes several minutes. Any tricks to speed it up?
2) Is there a better (or different) online forum for this software, other than this thread?
3) If a file has two sets of subs (same language), is there a way to work on one, then load up the other, without re-loading the file from scratch? Looking at your comment, "In the "Choose language" window you can do a "Save as..." for each language stream id." --> I didn't see any "save as" under "options-> choose language". Is that where you were talking about?
4) What does a red duration cell indicate?
5) My biggest, worst problem: I load the file, which takes several minutes, then I do OCR on PGS subs and spend a bunch of time fixing them all up, then I hit "OK" to leave the OCR dialog, and .... all my work is gone. The program doesn't crash. It's just that nothing at all shows up in the main window. So I had nothing to show for all my work. It happened to me several times, so I never got any complete srt out of it, except one time when I thought I was going to lose all my work yet again, I got it to write out an SRT with only a few subs as a test. Not sure what's going on. I'm afraid to put any work into editing subs, for it to all just go away when I click "OK". Please advise. (TL;DR: sometimes the OCR work is copied into the main window, but often not, and I haven't figured out why it behaves one way sometimes, and the other way sometimes.) This is a huge deal breaker for me, as I don't know if all of my work will be lost, or not. I think it might have something to do with click-and-dragging the mkv file (with subs) into the window to load it, vs loading using the "Open" dialog.
Thanks for the hard work!
Nikse555
20th April 2020, 07:23
@jlw_4049:
>Once minimized it cannot be re-opened until it's done. Which isn't a negative for me, but just letting you know!
If I click on the "Maximize icon" in the lower left corner, it comes back...
"Tools" - "Batch convert" can also OCR, but you'll loose the possible bad words etc.
@eddified:
"Binary image compare" takes a bit longer to get running (and learn). You must find the best "x pixels is space value" and you must add letters and fix wrong letters (double click on a line in the list view to inspect/fix). It does not work well for *all* subtitles, but it's especially nice if you got more than one file with the same font.
Tesseract is good too, very easy, but if problems occur they tend to be harder to fix. Also a little slower.
1) It takes about 1 second here for a 25 mb file...
2) This is probably the best forum for OCR stuff
3) That's for Vob ripping from DVD...
4) Red background color in duration cell indicates that the duration if too short or too long - see options - settings - general (min/max/cps)
5) Sorry about that, it's a bug in SE 3.5.14 with drag'n'drop (File - open, works fine)
And please test latest beta as SE 3.5.15 should be out soon :)
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
GCRaistlin
20th April 2020, 13:14
Nikse555
Bugs:
I use Ctrl-Win-X as a system-wide hotkey. Surprisingly pressing Ctrl-Win-X in SE calls 'Delete one line?' prompt window. I believe SE don't make the difference between Ctrl-X and Ctrl-Win-X. What is more I was unable to find how to disable Ctrl-X keyboard shortcut.
Trying to shift forward one subtitle so it would overlap the next subtitle is being corrected silently by SE. Example:
9
00:08:29,018 --> 00:08:30,810
<some text 1>
10
00:08:31,510 --> 00:08:34,018
<some text 2>
If we try to shift subtitle 9 for 800 ms forward then we get "00:08:31,653 --> 00:08:34,161" instead of the expected "00:08:32,310 --> 00:08:34,818".
Feature requests:
It's better to pre-select 'Selected line(s) and forward' instead of 'All lines' in Adjust all times... dialog window. If we want to shift everything we'll probably do it without selecting another line after opening the dialog window so 'Selected line(s) and forward' would mean the same as 'All lines' in most cases; if we select another line we probably don't want to shift everything so default 'All lines' has no sense in this case. 'All lines' option seems to me superfluous at all.
Ot at least remember the last used selection in this window.
Disable mouse wheel in 'Start time' and 'Duration' fields. Here's why. We click on Up or Down arrow in 'Duration' field to adjust the subtitle. Then we want to scroll the subtitle list with the mouse wheel. But instead of scrolling the subtitle list we get scrolling of the current subtitle duration!
Don't place the currently selected subtitle to the center of the screen when perform Undo/Redo action. It prevents us to see what we undo/redo.
Show the asterisk that indicates unsaved changes before the filename, not after. This way it will be always visible. Now it isn't if SE window isn't maximized or if there is an OSD of some other application in the top right corner of the screen.
Just in case you missed it: what you think of my suggestion #1 here (https://forum.doom9.org/showthread.php?p=1904263#post1904263)? It would give us great flexibility for subtitles adjusting. Now we have to perform many manual calculations, e. g. if we want to shift the subtitle start time without changing the subtitle end time.
GCRaistlin
20th April 2020, 15:12
'History (for undo)' window isn't intuitive. Example: I wanted to apply delay +200 ms to the selected subtitle and forward. I did it but I'm not sure if I have selected "Selected line(s) and forward". I go to Edit -> Show history (for undo) and see there: "Before show selected lines earlier/later: 00:00:00,200". It is completely unclear what does it mean. But okay. I press "Compare with current" and see that no, I didn't select the right option before applying the delay. I press 'Rollback'. The window closes; it is unexpected - I'd prefer it to stay open to give me the possibility to "Compare with current" again. But okay. For some reason I want to redo the action - I press Ctrl-Y. Then I press Ctrl-Z again. I perform these actions a couple of times as I'm sure that it doesnt' touch anything but the last action. But then I found out that I was wrong - now I can't undo anything that I've done before applying this delay!
That's how it should work according to my opinion:
Undoing the action isn't an action itself and should not overwrite other actions in the undo stack. It's the most important thing.
Actions in the history should be named natively. In my case above - "Delay +00:00:00,200 for all lines". If the delay was applied for the selected lines: "Delay +00:00:00,200 for lines ##10-14". And so on.
Actions that are undone should be displayed as Italic.
There should be two buttons in the dialog window: "Undo" and "Redo". If the selected action is undone the "Redo" button is active and "Undo" button is disabled. And vice versa.
If the selected action isn't the last action Undo/Redo are applied to the selected action and all actions after it.
GCRaistlin
20th April 2020, 20:58
The cursor in 'Start time' (in the main window) and 'Hour:min:sec:ms' (in 'Adjust all times...' dialog window) fields is in Overwrite mode: what we enter overwrites the current value. The cursor in 'Duration' field is in Insert mode: what we enter doesn't overwrite the current value. It would be better if the cursor was in Overwrite mode everywhere.
Nikse555
21st April 2020, 15:48
@GCRaistlin: thx for the feedback :)
1) About "Ctrl-Win-X" hotkey - I presume that SE should ignore all shortcuts if a Win-key is down, right?
2) Undo... I think what I actually wanted to show is more an "Event log" - where *all* entries get saved and can be rolled back to. I am really annoyed by the undo in many gfx application where when you undo something that you want to re-do later, but suddenly the history has been cleared due to a new change and you cannot re-do as you planned (and changes are lost)! I can see that the current "History for undo" is kinda stuck in the middle...
3) Duration field overwrite mode... yes, please. But I actually don't know how.
4) I can see the idea in the asterisk that indicates unsaved changes before the filename, not after...
GCRaistlin
21st April 2020, 17:10
1) About "Ctrl-Win-X" hotkey - I presume that SE should ignore all shortcuts if a Win-key is down, right?
Yes. Besides that I believe that all keyboard shortcuts should be customizable (since you have implemented this for some of them anyway).
I am really annoyed by the undo in many gfx application where when you undo something that you want to re-do later, but suddenly the history has been cleared due to a new change and you cannot re-do as you planned (and changes are lost)!
This maybe makes sense but repeatedly undoing/redoing kills the possibility of undoing previous actions this way. I believe it is worse than a killing new change because it is completely unexpected.
3) Duration field overwrite mode... yes, please. But I actually don't know how.
You did it for 'Start time' field but you don't know how to do it for 'Duration' field?
Can you please make 'Adjust all times' dialog window auto-closing on 'Show earlier' or 'Show later' press?
Nikse555
21st April 2020, 18:02
Yes. Besides that I believe that all keyboard shortcuts should be customizable (since you have implemented this for some of them anyway).
Yes, not many hardcoded shortcuts left now :)
Latest beta should ignore shortcuts when the Windows-Key is pressed:
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
Does that work for you?
You did it for 'Start time' field but you don't know how to do it for 'Duration' field?
Yeah, sorry. Winforms does not have the best controls in the world. It's of course possible in some way but there's no quick fix to do this for a "NumericUpDownControl" as far as I know.
Can you please make 'Adjust all times' dialog window auto-closing on 'Show earlier' or 'Show later' press?
I often keep the window open or click "Show later" multiple times...
Nikse555
22nd April 2020, 06:49
Titlebar now also changed so asterisk is first: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
tormento
26th April 2020, 09:05
Titlebar now also changed so asterisk is first
Hi Nikse!
One regression (to me) was the "fix continuation style". I was really happy with the removal of leading … and now that new rule is making sometimes a mess.
Some requests:
in the binary compare dialogue, right clicking in the input field gives some used characters in different language families. You forgot italian: we use same vowels as spanish but with open accent, i.e. àèìòù and é plus all their capitals. For the lower case it's easy, as we have on keyboard, for the upper I have to use charmap. And please add french letters too, such as âêîôû and capitals. Perhaps it's better to do a separate menu for vowels and variations only, as they repeat in a lot of languages.
set a new rule to find mixed case words, usually OCR errors, such as RAvEN or raVen, would you?
the possibility to order the rule order in the fix errors dialogue or preferences
move the red big Italic warning to the right of the character to type, it's more visible there
add a green Regular warning too
add a flag not to add the typed character in the OCR DB: sometimes there are strange ones that would cause issues on later OCR
report the OCR progress bar on the application bar Subtitle Edit icon (such as other app do) and/or ring a bell when it needs attention
Thanks!
Nikse555
26th April 2020, 15:03
Hi tormento!
I'm just finishing up SE 3.5.15 and only bug fixes right now - but I will keep you input in mind.
The "Continuation style" is now default "not-enabled" (works like it used to be). You can enable (change style or disable) it via Options -> Setting -> General - and "Continuation style" in profile.
The SE OCR taskbar icon will now blink (for a while) when input is required and SE is not focused. Is that what you mean by #7?
Last chance to find bugs before 3.5.15 final :)
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.14/SubtitleEditBeta.zip
tormento
26th April 2020, 15:11
The "Continuation style" is now default "not-enabled"
Nice and I would like to see "remove leading ..." back :)
Is that what you mean by #7?
Yep and I'd like some bells too. :)
Last chance to find bugs before 3.5.15 final :)
There is a really strange bug when smart fixing italian subs that transform every I or l in L. I reported few posts ago. It's rare but sometimes it happens. I will check it out to see if it happens again and give you some sub to test.
jlw_4049
27th April 2020, 05:35
Does the batch tool use tesseract? There is no way to select which one it uses. So I am unsure as to which is being used.
Nikse555
27th April 2020, 06:53
@tormento: "Rremove leading ..." is back when not using "Continuation style".
Bells... at first blink only or at every blink? or ?
@jlw_4049: Yes, the batch tool uses Tesseract, with last used language.
tormento
27th April 2020, 08:15
@tormento: "Rremove leading ..." is back when not using "Continuation style". Bells... at first blink only or at every blink? or ?
Thanks, you could move remove leading to continuation style in options.
For every, let’s say, 5 blinks. If you have time put a value in options, with enabling or disabling sound.
Janusz
28th April 2020, 09:35
Hello.
Sorry, I don't know English so I used a translator.
Many thanks to the author for this excellent program and to everyone who develops it.
I use it for many years mainly to extract subtitles from the * .ts stream broadcast by TV stations. I use nOCR because I think this method is unrivaled in terms of speed and accuracy.
Now my comments:
1. With one station it is such that there is not enough space between the first and second line of subtitles and OCR cannot deal with it.
https://forum.doom9.org/attachment.php?attachmentid=17310&stc=1&d=1588067414
https://forum.doom9.org/attachment.php?attachmentid=17311&stc=1&d=1588067512
2. ',' (comma) normally and ',' (comma) used as ' (accent) in e.g. English spelling. I bypassed this problem by combining ' with a letter, but such a combination of characters by "expand" e.g. <' a> - you can't see the nocr character database. You also don't see, for example, % ("o" extended by another two characters "/ o"). Characters created by "expand" are lost when exporting / importing the character base.
3. <No of pixels is space> in <OCR method>
- If I import images from the * .ts file <No of pixels is space = 4>, (this is the configuration file <Settings.xml> and this value is saved for future reference.
- If I import images from index.html - <No of pixels is space = 12> - always,
no matter that I changed this value before.
- <No of pixels is space> for italics. In this case, it would be useful to use a different space than, for example, 4, and for italics 3. Decreasing the space by 1 for italics generates much less errors in the combined words. Especially for <j> in italics after <w, r, A>
Best regards, Janusz.
GCRaistlin
28th April 2020, 13:36
Latest beta should ignore shortcuts when the Windows-Key is pressed:
Does that work for you?
Yes, it does, thanks.
I often keep the window open or click "Show later" multiple times...
You can implement it as an option. At my opinion it's handier to close the dialog and then open it again by a keyboard shortcut.
Also, it would be nice to have the following things implemented here:
Time input field is active after opening the dialog.
Underscored letters for quick access the radio buttons and keys on Alt-<letter>: Show earlier, Show later, Selected line(s) only, Selected line(s) and forward, All lines.
This way one can call the dialog and have access to all its controls with the keyboard and, what is more, by using only his left hand (except 'All lines' but it doesn't really matter).
Is there a way to expand the selection forward when performing OCR? The first part of Cyrillic "ы" coincides with Cyrillic "ь".
Found a solution: delete "ь" from the database, define "ы", define "ь" again.
It isn't the best solution though as "ь" could be defined when performing OCR of a different sup file so it may not be easy to redefine "ь" again. Can you please add the possibility to add a better match with expanding the selection?
It is unable to load a DVD sup file from within UI while it is possible supplying it as a command line argument.
GCRaistlin
30th April 2020, 13:10
Feature requests:
Check the file being open for errors (erroneous example (https://mir.cr/6FD66P54)). DVDSubEdit does.
'Skip current subpic' button in 'Manual image to text' OCR dialog. It will allow not to abort OCR if the scanned file contains errors (erroneous example (https://mir.cr/018C3DT1), see #358).
GCRaistlin
30th April 2020, 14:15
The issue that is similar to what I reported earlier (https://forum.doom9.org/showthread.php?p=1904263#post1904263) (# 2). The same example as for FR # 2 above. Try to OCR it with my Cyrillic.db (https://mir.cr/12QAJSAB). If you start from the beginning then the 2nd word in the 2nd line of # 231 will be recognized as "остаьаться". If you start from # 231 itself it'll be "оставаться".
What is interesting is that in 'Inspect compare matches for current image' dialog the problematic character seems to be recognized properly:
https://i111.fastpic.ru/thumb/2020/0430/73/b3adfdb2898006944048a7f8ef9c4073.jpeg (https://fastpic.ru/view/111/2020/0430/b3adfdb2898006944048a7f8ef9c4073.jpg.html)
darksen
30th April 2020, 20:59
SE 3.5.12 is out :)
@darksen: Did you get a lot of false corrections? Perhaps that could be improved if you posted some.
Sorry for taking so damn long to reply back.
Yes, I always get a lot of that. I make subtitles in Spanish and obviously I use the Spanish rules, some errors I get very often are about the OCR correction, sometimes, and this is just one example, SE wants to replace Marlboro with Mariboro (l for i), other times it wants to replace names for their localized ones, like Ivan with Iván, other times it wants to fix quotes when it doesn't have to, in Spanish the main quotes we use are angled quotation marks (« and ») and if there is a quote inside that quote we use double quotes (" and ") and if there is another quote inside that we use single quotes (' and ' ) but SE sometimes wants to fix the double quotes and replace them with the angled ones despite they being inside angled already, i.e:
Ivan said: «Blahblahblah "blah-blah-blah" blahblah»
SE wants to fix it like:
Ivan said: «Blahblahblah blah-blah-blah blahblah»
Or sometimes:
«Blahblahblah «blah-blah-blah» blahblah»
It also happens when the dialogue is split in more than one subtitle because SE doesn't detect that angled quotes have been used in the previous sub.
And there have been more errors but I can't remember them all right now. That's why I'm asking to split the "Fix common OCR errors (using OCR replace list)" into two, one informing that the default list is used for said correction and the other that the correction comes from the user list.
Thank you.
GCRaistlin
30th April 2020, 22:34
Feature request: split 'No of pixels is space' ('Import/OCR Blu-ray (.sup) subtitle' dialog window) to 'No of pixels is space after an Italic character' and 'No of pixels is space after an non-Italic character'. These values should definitely differ.
GCRaistlin
30th April 2020, 23:35
Feature requests:
Quick paste of characters with acutes and umlauts in 'Inspect compare matches for current image' dialog (as in 'Manual image to text' dialog).
The possibility to paste a subtitle from the clipboard. Here's what I mean. Let's consider I have two srt files open in two different copies of SE. In one copy, I do a right-click on a line and select 'Copy as text to clipboard'. In another copy, I do a right-click on a line and select 'Paste from clipboard before'. SE adds a subtitle from the clipboard preserving everything but subtitle number (i. e. start time, duration, text). Why do I need this? It may be useful when syncing subtitles (*) that we downloaded from some site to the ones (**) that we ripped from the disc we want to view with (*) and when subtitles (**) are lacking of some lines that are present in subtitles (*). This way, we could easily copy-and-paste missing lines from (*) to (**), then shift them by applying a delay, then import time codes from (**) to (*).
Janusz
1st May 2020, 12:23
Feature request: split 'No of pixels is space' ('Import/OCR Blu-ray (.sup) subtitle' dialog window) to 'No of pixels is space after an Italic character' and 'No of pixels is space after an non-Italic character'. These values should definitely differ.
That's exactly what I mean, as I wrote above.
Feature requests:
Quick paste of characters with acutes and umlauts in 'Inspect compare matches for current image' dialog (as in 'Manual image to text' dialog).
The possibility to paste a subtitle from the clipboard. Here's what I mean. Let's consider I have two srt files open in two different copies of SE. In one copy, I do a right-click on a line and select 'Copy as text to clipboard'. In another copy, I do a right-click on a line and select 'Paste from clipboard before'. SE adds a subtitle from the clipboard preserving everything but subtitle number (i. e. start time, duration, text). Why do I need this? It may be useful when syncing subtitles (*) that we downloaded from some site to the ones (**) that we ripped from the disc we want to view with (*) and when subtitles (**) are lacking of some lines that are present in subtitles (*). This way, we could easily copy-and-paste missing lines from (*) to (**), then shift them by applying a delay, then import time codes from (**) to (*).
Instead of starting two sessions of the program, did you try "translator mode"
with the second file as "original text"
Whereby: <Opion / Settings / Allow edit of original subtitle> = "enabled"
I would need a new function:
If I mark a line in the "Compare subtitle" window, it is automatically selected in the main program window.
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.15
Janusz
1st May 2020, 19:41
@GCRaistlin: 2. The possibility to paste a subtitle from the clipboard. Here's what I mean. Let's consider I have two srt files open in two different copies of SE. In one copy, I do a right-click on a line and select 'Copy as text to clipboard'. In another copy, I do a right-click on a line and select 'Paste from clipboard before'. SE adds a subtitle from the clipboard preserving everything but subtitle number (i. e. start time, duration, text). Why do I need this? It may be useful when syncing subtitles (*) that we downloaded from some site to the ones (**) that we ripped from the disc we want to view with (*) and when subtitles (**) are lacking of some lines that are present in subtitles (*). This way, we could easily copy-and-paste missing lines from (*) to (**), then shift them by applying a delay, then import time codes from (**) to (*).
You have tried: Right click on the line and select <Column / Paste from clipboard ...>
There is what you wrote about.
GCRaistlin
1st May 2020, 22:00
Janusz
It is pretty good hidden. Thanks.
locotus
2nd May 2020, 18:52
Going back to 3.5.14 after noticing that remove leading (...) was remove. Sorry for that.
Nikse555
2nd May 2020, 19:08
@locotus: "Remove leading ..." is still there with default settings: https://nikse.dk/fix-common-errors.png
But it depends on what "Continuation style" is chosen in "Profile". If you like dash/epllises for both start/end then "Remove leading ..." does not make sense. Read more here: https://github.com/SubtitleEdit/subtitleedit/pull/4108
Thanks for new release:
https://github.com/SubtitleEdit/subtitleedit/releases/tag/3.5.15
Your're welcome.
Yes, 3.5.15 is out :)
Change log: https://raw.githubusercontent.com/SubtitleEdit/subtitleedit/master/Changelog.txt
@GCRaistlin: thx for the sup file, SE should now allow some errors in the file (like for .ts files).
You can also do File -> Import time codes...
locotus
2nd May 2020, 19:24
@locotus: "Remove leading ..." is still there with default settings: https://nikse.dk/fix-common-errors.png
But it depends on what "Continuation style" is chosen in "Profile". If you like dash/epllises for both start/end then "Remove leading ..." does not make sense. Read more here: https://github.com/SubtitleEdit/subtitleedit/pull/4108
All of the contrary, I only use ellipsis at the end of a dialog that needs to be continue, but never at the beginning
of the continuation because it means losing 3 characters at the start of an already long line and this is what don't make any sense.
Is there any continuation style that only use ellipsis at the end of the line?
Nikse555
2nd May 2020, 19:35
Is there any continuation style that only use ellipsis at the end of the line?
You can try "Dots (trailing only)": This will also add "..." for continuing lines where it's not present.
There's also "None, dots for pauses (trailing only)".
locotus
2nd May 2020, 19:55
You can try "Dots (trailing only)": This will also add "..." for continuing lines where it's not present.
There's also "None, dots for pauses (trailing only)".
Thanks for your time but I'd rather stay with 3.5.14.
tormento
3rd May 2020, 10:25
Yes, 3.5.15 is out
There is a nasty OCR correction in Italian "Fix OCR errors":
It's unusual a «I» is inside a period but when it's the initial letter of a name.
I.e. «I'» is impossibile in italian language, it should always be «l'».
Likewise, «I*» where * is another letter, is very difficult to be found, unless it is the beginning of a personal name or an all capital letter word.
I suggest you to replace every «I» with «l» unless inside a word or at the beginning of a period or paragraph.
The only italian word of two letters that begins with a «I» is «Io», i.e. «I» in english. There are instead definite article «lo», i.e. "male" «the» in english,.usually contracted as «l'» in front of a vowel, that can be a personal pronoun too, i.e. «him» in english. The latter can be contracted as «l'» too.
Nikse555
3rd May 2020, 10:40
@tormento: Could you give one or two examples? How does it work now and how should it work?
Also, Italian context-menu with letters should now be included in the "OCR character" + "OCR inspect" windows: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
tormento
3rd May 2020, 11:30
@tormento: Could you give one or two examples? How does it work now and how should it work?
Of course, ASAP I will find some proper ones.
"OCR character" + "OCR inspect" windows
Where is it? I can't find it.
P.S: I suggest you to exclude from merge lines the ones with capital words inside. Include in the list but uncheck them as default.
tormento
3rd May 2020, 13:04
First example of not corrected «l»:
893
01:28:03,875 --> 01:28:06,500
IL VERO HACHIKO NACQUE
A ODATE, lN GIAPPONE, NEL 1923.
894
01:28:06,583 --> 01:28:08,750
QUANDO lL SUO PADRONE,
IL DOTT. EISABURO UENO,
896
01:28:11,875 --> 01:28:13,667
HACHI TORNÒ lL GIORNO DOPO AD ASPETTARLO
897
01:28:13,750 --> 01:28:16,583
ALLA STAZIONE DI SHIBUYA E
COSÌ FECE PER l SUCCESSIVI 9 ANNI
900
01:28:28,708 --> 01:28:31,667
UNA STATUA lN BRONZO DI HACHIKO
SI TROVA NEL PUNTO ESATTO
901
01:28:31,750 --> 01:28:33,708
IN CUI lL CANE ASPETTAVA lL SUO PADRONE
It should be:
893
01:28:03,875 --> 01:28:06,500
IL VERO HACHIKO NACQUE
A ODATE, IN GIAPPONE, NEL 1923.
894
01:28:06,583 --> 01:28:08,750
QUANDO IL SUO PADRONE,
IL DOTT. EISABURO UENO,
896
01:28:11,875 --> 01:28:13,667
HACHI TORNÒ IL GIORNO DOPO AD ASPETTARLO
897
01:28:13,750 --> 01:28:16,583
ALLA STAZIONE DI SHIBUYA E
COSÌ FECE PER I SUCCESSIVI 9 ANNI
900
01:28:28,708 --> 01:28:31,667
UNA STATUA IN BRONZO DI HACHIKO
SI TROVA NEL PUNTO ESATTO
901
01:28:31,750 --> 01:28:33,708
IN CUI IL CANE ASPETTAVA lL SUO PADRONE
tormento
3rd May 2020, 14:47
This (https://www.mediafire.com/file/cm2scqov6h1i2ai/The_raid_sup.7z/file) sup gives me problems with Binary image compare.
It seems it is changing color palette and SE can't keep the pace to it. Plus there is graphic corruption when palette change is encountered.
tormento
3rd May 2020, 15:30
When binary compare OCR subtitles with italic, there are some issues with spaces between words.
Here (https://www.mediafire.com/file/ayayqow1oizpeu1/The_spy_who_sup.7z/file) an example, where it's impossible to have a proper space for normal text and italic on line 9 (ofpower).
The binary compare should apply 2 different "number of pixels is space" to italic and normal text.
Nikse555
3rd May 2020, 15:46
SE is relying on the "???_OCRFixReplaceList.xml" file to make corrections... sadly there's no "ita_OCRFixReplaceList.xml". I've tried to make one in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Your examples above seem to be fixed I think, but I'm not sure if it breaks something else...
Also, perhaps you could add some fixes/words to it? Perhaps you already have a "ita_OCRFixReplaceList_user.xml" file? You can take a look at the "eng_OCRFixReplaceList.xml" for inspiration.
About the palette, a bit tough, but if you click on "image pre-processing", and select "100" for thresshold, color2white: choose default black, color2remove: choose default white...
Some subtitles uses transparent for "color"... SE cannot handle those.
SE already tries "number-of-pixels" -1 for italic, but I will take a look at the file. Also a lot can be fixed via the OCR fix replace list.
Also, above beta you can customize the "quick-click" letters in the "OCR character" window - search for "OcrAddLetterRow" in Settings.xml. You can also right-click in the text box and get a context menu where Italian letters are available.
Nikse555
3rd May 2020, 16:00
When binary compare OCR subtitles with italic, there are some issues with spaces between words.
Here (https://www.mediafire.com/file/ayayqow1oizpeu1/The_spy_who_sup.7z/file) an example, where it's impossible to have a proper space for normal text and italic on line 9 (ofpower).
The binary compare should apply 2 different "number of pixels is space" to italic and normal text.
Works fine here with a clean default installation with "pixels is space" set to 9 (edit: up to 11 works fine).
tormento
3rd May 2020, 16:37
Also, perhaps you could add some fixes/words to it?
I will do my best in spare time.
About the palette, a bit tough
Had a hard time finding where image preprocessing was but now it works, thanks!
Works fine here with a clean default installation with "pixels is space" set to 9 (edit: up to 11 works fine).
Really strange. What setting could influence mine not working?
Janusz
3rd May 2020, 17:09
This (https://www.mediafire.com/file/cm2scqov6h1i2ai/The_raid_sup.7z/file) sup gives me problems with Binary image compare.
It seems it is changing color palette and SE can't keep the pace to it. Plus there is graphic corruption when palette change is encountered.
First, I did <Save all images with HTML index ...>
Then I did OCR from the index.html file.
The program has read all lines, including those where the colors change.
GCRaistlin
3rd May 2020, 19:51
Nikse555
Bugs:
If SUP file is open by supplying it as an argument in the command line and got OCRed then after loading the OCR results to the main window SE behaves as if an unchanged SRT file is open: no asterisk in the window title, no prompt to save file on closing. If we won't save the file manually we'll lose our work.
The lack of a point isn't highlighted in 'Compare subtitles' - try to compare these ones (https://mir.cr/0T8EVLL9).
Open a BD SUP file (from the UI), go to, say, 1000th subpic, press 'Start OCR' - 'Stop' - 'Cancel'. SE prompts if we want to discard changes, as expected.
Open the same file (from the UI), don't go anyway, press 'Start OCR' - 'Stop' - 'Cancel'. SE closes without any prompts - NOT as expected.
Open a DVD SUP file (by supplying it as an argument in the command line), go to, say, 1000th subpic, press 'Start OCR' - 'Stop' - 'Cancel'. SE closes without any prompts - NOT as expected.
When we export subpic # 1078 as an image SE offer Image1077.png as a filename.
Open an ASS file (https://mir.cr/AZ3FGZF5). Do 'Copy as text to clipboard' on any line, then 'Column' - 'Paste from clipboard...'. You'll get garbage pasted.
Feature requests:
The ability to define keyboard shortcut for <right-click> - Column - Paste from clipboard...
Make 'Column paste' dialog remember last used radio buttons states. Also, add underscored letters to the radio buttons to make it possible to select it with Alt-<letter>.
The ability to load DVD SUP files from UI.
Why bitmaps that was added to OCR DB during performing OCR of a SUP file don't work when performing OCR of PNG files exported from this SUP file (by 'Save subtitle image as...')? Try to OCR 00262.track_4608.sup (https://mir.cr/0AAWJH75) using my Latin.db (https://mir.cr/0CCIMJS4) from # 1078 - SE won't ask you anything. Then export subpic # 1078, import it and try to OCR it.
You are welcome to import new records from my Latin.db to the supplied one.
Janusz
3rd May 2020, 20:55
...
Feature requests:
1. The ability to define keyboard shortcut for <right-click> - Column - Paste from clipboard...
Use <Options / Settings / Shortcuts>
In the Search box, type: Col
at the bottom you can also choose "Column, paste". Set the shortcut you want.
Nikse555
3rd May 2020, 21:20
@GCRaistlin:
>1. After loading the OCR results to the main window SE behaves as if an unchanged SRT file is open
I cannot re-create this... what file format and how do you open the file exactly?
>3. Open a BD SUP...
Yeah, that's a feature. Do any real work and SE will prompt.
>3. The ability to load DVD SUP files from UI...
That works here in main window via File - Open... or drag-n-drop. How/where does it not work?
GCRaistlin
3rd May 2020, 21:40
Janusz
Thanks again. I saw it but didn't think that this is what I'm looking for. Just wondering why it has a different name here...
Janusz
3rd May 2020, 22:48
@GCRaistlin
Bugs:
1. After loading the OCR results to the main window SE behaves as if an unchanged SRT file is open: no asterisk in the window title, no prompt to save file on closing. If we won't save the file manually we'll lose our work.
1. I don't know how you are.
For me in the title of the main window after loading the subtitles after OCR the window name changes to: * D: \ full path \ filename.html \ index.srt - SubtitleEdit...
If I am now trying to close the program I get a prompt to save a new file. I don't know why it doesn't work for you.
After saving, the window name indicates where to save the index.srt file without *. If I change anything in the text, the program name will start with * again.
GCRaistlin
3rd May 2020, 23:26
Janusz
I got it: it happens when SUP file is open by supplying it as an argument in command line, not from UI.
Nikse555
4th May 2020, 12:37
@GCRaistlin: Beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Fixed dvd sup from cmd line + ass paste + change detection with file+ocr from cmd line + compare issue + image save as number + remember column paste options - hopefully some of that also works you?
Janusz
4th May 2020, 17:33
@Nikse555
I don't know how understandable this translator text will be, but I'll try.
In connection with the problem of correct recognition by the OCR systems of the lowercase "L" and the "I",
please explain briefly the principles which follow Subtitle Edit for automatic correction during OCR.
Why i ask?
I start the OCR process with a new, completely empty character base.
My settings: in [OCR auto correction Dictionary] - none, other options unchecked.
No dictionaries in the Dictionaries catalog. Options/Tools [Fix common OCR errors ...] unchecked.
As a result, I get an empty text with spaces recognized according to <No of pixels is space>.
This is correct and as expected.
I add the lower case letter "L" to the character base, options marked as in (1).
As a result, they appear, I don't know if all but definitely recognized lower case "L", but also "I",
which I don't have in the character database. Spaces between words as in (1).
I can agree. At this stage I accept "I", because in Polish lonely "I" occurs so it is OK.
But why does the program change the lowercase "L" to the "I" at the beginning of words
since they do not yet know these words. It looks like this: [I I**** *********].
The first "I" - ok. The second may well be the lower case letter "L".
I consider this a serious mistake.
Such a conversion could take place only after recognizing the entire word and after checking
in the selected dictionary that such a conversion would not cause an error.
From what you can see - checking in the dictionary is missing or not working as it should.
I omit the fact that now the dictionary is off because every word is a mistake.
Another OCR attempt.
a. I remove the only lower case letter "L" from the character base. Checking - the character database is empty. Start OCR - result as in (1) - OK.
b. Another OCR attempt. The character base is the same. However, the small "L", previously removed, is still there. I do not know why?
Did the program not save the changes permanently? OCR result as for (2) with all errors ("I").
I tried in various ways to get rid of the stubborn letter from the base,
but you can't delete a character if it is the only character in the database.
The changes only apply in the current program session.
After closing and restarting the program, everything (last letter) returns to the state it was before.
I think it's a mistake. Or maybe it should work like that? I do not know.
Finally, my suggestion to consider in the distant or near future.
When automating the OCR process, give the opportunity to use (set) a second dictionary. Subtitles are usually a translation of one language into another.
However, proper names, first names, last names etc. which do not have equivalents in a given language are often not translated.
Using a second language outside the main language - will generate fewer errors, making the process more intelligent.
We don't always create or correct subtitles in one language.
If the translation is not understandable enough, I am sorry.
I wanted to help You and myself.
tormento
4th May 2020, 20:52
@GCRaistlin: Beta updated
I really can't find a pixel space to binary OCR the italic part of this (https://www.mediafire.com/file/hx912s8z3ftw56b/500_sup.7z/file) sup. No problems at all with SupRip.
Can you please tell me if you can and what parameters are you using?
GCRaistlin
5th May 2020, 00:28
Fixed dvd sup from cmd line
Not sure what you mean. The other fixes are confirmed, thank you.
>3. Open a BD SUP...
Yeah, that's a feature. Do any real work and SE will prompt.
I did the real work actually: I opened a BD SUP file, started OCR from the beginning, then aborted it at some point. There is recognized data, but if I press Cancel SE discards it silently. Though if I do the same steps but start OCR from # 1000 SE asks me about discarding. Why the behavior is different?
My logic is pretty simple: if there is any recognized data (even just one char) SE should ask about discarding if the user press Cancel or tries to close the window.
>3. The ability to load DVD SUP files from UI...
That works here in main window via File - Open... or drag-n-drop. How/where does it not work?
Yeah, this works. I should say though that it is completely unintuitive: we should use 'File - Import/OCR' for BD SUP files but 'File - Open' for DVD SUP files.
Bug: if we open a DVD SUP from the command line and then press Cancel in 'Import/OCR' dialog SE remains open.
Nikse555
5th May 2020, 11:47
@GCRaistlin:
"Fixed dvd sup from cmd line" is about the "*" in the title bar.
OK, OCR window should now prompt for save changes if anything has been added.
File - Import/OCR' for BD SUP will now also allow dvd sup. I just always use File -> Open...
>Bug: if we open a DVD SUP from the command line and then press Cancel in 'Import/OCR' dialog SE remains open.
Thx, should now hopefully be fixed.
Latest beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
I really can't find a pixel space to binary OCR the italic part of this (https://www.mediafire.com/file/hx912s8z3ftw56b/500_sup.7z/file) sup. No problems at all with SupRip.
Can you please tell me if you can and what parameters are you using?
Sorry, SE cannot handle that file via "Binary image compare".
tormento
5th May 2020, 12:05
Sorry, SE cannot handle that file via "Binary image compare".
SupRip tilt OCR frames according to character vertical lines.
Isn't possible to implement something like that into SE?
tormento
5th May 2020, 13:47
@Nikse
This (https://www.mediafire.com/file/xtf28xmnndnc8bf/apes_sup.7z/file) too has problem with Italic. Perhaps there was some regression at some point, because almost all titles I am doing OCR are "cursed".
Can you fix the binary compare engine, regarding Italic? You are telling that it's easily fixed by OCR error rules but I can't find any universally suitable.
GCRaistlin
5th May 2020, 15:50
File - Import/OCR' for BD SUP will now also allow dvd sup. I just always use File -> Open...
Maybe it's better then to remove 'Import/OCR VobSub' and 'Import/OCR BluRay (.sup)' menu items since such files can be opened via File - Open? If you think it's too much then please at least rename 'Import/OCR BluRay (.sup)' to simple 'Import/OCR .sup subtitle file'.
Feature requests:
Ability to set font properties for Spell Check dialog window.
Ability to run Spell Checking from within 'Import/OCR' window after performing OCR - with the possibility to select the subpic with the currently unknown word. Why don't I perform spell checking during the OCR? Because I prefer adding a better match than using 'Fix common OCR errors' or spell check (it's better to prevent errors than to fix them). I consider spell check as a tool that helps finding errors that were missed first during the OCR and then during visual check.
A lot of mistakes in word boundaries detection could be avoided if the vertical lines of a character, if any, were taken into account in the first place. Therefore we need a new setting - 'No of pixels from/to vertical line is space'; this setting takes precedence over simple 'No of pixels is space'. Examples:
"of July" in this subpic:
https://i111.fastpic.ru/thumb/2020/0505/71/71c4ab87d968ccea6422ddf2ccdb0871.jpeg (https://fastpic.ru/view/111/2020/0505/71c4ab87d968ccea6422ddf2ccdb0871.png.html)
is now being recognized as "ofJ uly" - two errors in once (with 'No of pixel is space' setting of 9 which seems to be the most reliable choice for all subpics in the sup file). The modified algorithm could process it as follows:
For "f", there is 27 running dots with the same X position from the right boundary of 47 total "vertical" dots; hence, we have a vertical line from the right side here.
For "J", there is 46 running dots with the same X position from the left boundary of 56 total "vertical" dots; hence, we have a vertical line from the left side here.
There are 26 pixels between these two vertical lines while there are only 6 pixels between the most right dot of "f" and the most left dot of "J".
There are 10 pixels between "J" and "u".
"If you" in this subpic:
https://i111.fastpic.ru/thumb/2020/0505/41/9f8eb4b995d0e5d8e8ba83ccb473c041.jpeg (https://fastpic.ru/view/111/2020/0505/9f8eb4b995d0e5d8e8ba83ccb473c041.png.html)
is now being recognized as "Ifyou" (with 'No of pixel is space' setting of 9). The modified algorithm could process it as follows:
We have the same case with "f" as above.
With "y", we don't have a vertical line from the left side as there's no enough quantity of running dots with the same X position from the left boundary.
Hence, we take into account the right vertical line of "f" and the most left dot of "y" when counting pixels between these characters: 19 pixels.
With 'No of pixels from/to vertical line is space' setting of 19 and 'No of pixel is space' setting of 10 both problematic substrings above would be recognized properly.
GCRaistlin
5th May 2020, 22:24
Feature request: use different fonts for list view and text boxes. For text boxes, Courier New is sometimes a better choice than Tahoma as it clearly shows the difference between a double quote and doubled apostrophe (OCR error). But list view looks ugly with Courier.
GCRaistlin
6th May 2020, 01:22
There's some incompatibility with Ditto (https://sourceforge.net/projects/ditto-cp/files/?source=navbar), the clipboard manager:
Download and install Ditto.
Import the .reg file:
[HKEY_CURRENT_USER\Software\Ditto]
"DittoHotKey"=dword:00000857
It sets Win+W keyboard shortcut to call Ditto.
Run Ditto, copy some text to the clipboard.
Open SE, place the cursor to the text box.
Press Win+W, select the text in the list, press Enter.
The text is pasted. Now you are unable to edit the text in the text box (though you can paste there).
tormento
6th May 2020, 11:10
If possible, do no automatically select remove line break when the next line starts with a capital letter.
varekai
6th May 2020, 16:53
Just wanted to say thanks for Subtitle Edit, couldn't do without it!
I'm not much of a bugtester (there seems to be others...;) as SE works perfectly for me!
So... I feel there's no need for changing the GUI.
varekai
6th May 2020, 16:56
@GCRaistlin
Please, consider another image hosting...:devil:
GCRaistlin
6th May 2020, 18:30
varekai
You are welcome to offer another one if it is as handy as the current one is.
varekai
6th May 2020, 19:03
This is how I would do it... (https://i.imgur.com/pmU53W7.jpg)
Nikse555
6th May 2020, 19:57
@GCRaistlin: Regarding Ditto, I have seen the error like 1 time in 100 pastes and I'm afraid I've no idea what's wrong or how to fix it. Ideas/fixes are welcome.
@tormento: Writing a new image-to-letter-splitter and integrating it into SE is a lot of work - could easily take 14+ days full time.
Is SupRip open source?
Edit: Thx for the sample files with italic! Surely uses more tilt than I've seen before.
Nikse555
6th May 2020, 20:01
Just wanted to say thanks for Subtitle Edit, couldn't do without it!
You're welcome, and thx :)
GCRaistlin
6th May 2020, 22:24
Regarding Ditto, I have seen the error like 1 time in 100 pastes
That's strange. I have the lock error every time I'm pasting from the Ditto list.
GCRaistlin
6th May 2020, 22:40
This is how I would do it... (https://i.imgur.com/pmU53W7.jpg)
Without the preview? A good change, indeed.
jlw_4049
7th May 2020, 03:29
A couple bugs. 'I' wants to be |
When using the latest version, both the original Tesseract 3x and 5x on certain parts of the PGS subtitle file is zoomed in to far for me or the program to read what it is. Other then those problems the newer version seems to perform better.
varekai
7th May 2020, 07:17
Without the preview? A good change, indeed.
(drumroll)
https://i.imgur.com/3DcrNTy.png
OKEYYY!
https://i.imgur.com/NeOY2LN.jpg
Are you just playing stupid or are you really... ;)
Edit
...clarification...
https://i111.fastpic.ru/thumb/2020/0505/71/71c4ab87d968ccea6422ddf2ccdb0871.jpeg
https://i111.fastpic.ru/thumb/2020/0505/41/9f8eb4b995d0e5d8e8ba83ccb473c041.jpeg
Nikse555
7th May 2020, 07:27
A couple bugs. 'I' wants to be |
Could you upload a sample where this happens?
SE tries to fix this via the eng_OcrFixReplaceList.xml file...
jlw_4049
7th May 2020, 07:52
Could you upload a sample where this happens?
SE tries to fix this via the eng_OcrFixReplaceList.xml file...I'll send the sample first thing tomorrow when I get up with the line(s) that it does it on.
Thanks!
Edit: Also adding the OCR corrected words in this format fixed the I issue.
-| to I
Sent from my Pixel 3a using Tapatalk
GCRaistlin
7th May 2020, 11:47
Nikse555
Regarding Ditto again. The issue seems to be worse that I thought: it is enough to just call the Ditto list on the screen (Win+W) having SE active to get the lock error. But it is only the main screen text box that gets locked - text boxes in Options and 'Import/OCR' window are immune to this.
Nikse555
7th May 2020, 15:28
Nikse555
Regarding Ditto again. The issue seems to be worse that I thought: it is enough to just call the Ditto list on the screen (Win+W) having SE active to get the lock error. But it is only the main screen text box that gets locked - text boxes in Options and 'Import/OCR' window are immune to this.
OK, how is latest beta?
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
GCRaistlin
7th May 2020, 17:35
Nikse555
Seems to be fixed, thanks!
jlw_4049
7th May 2020, 22:38
Could you upload a sample where this happens?
SE tries to fix this via the eng_OcrFixReplaceList.xml file...
So I re-opened the .sup today and it wasn't doing it at all. I walked a way for a bit after minimizing it and the display screen was enlarged when I opened it back up.
It was zoomed in about 50% to much and focused on the left side.
Here is the .sup I was able to reproduce it with.
http://www.mediafire.com/file/r69v2z6h7vyhkmk/sample.sup
Nikse555
8th May 2020, 10:06
...after minimizing it and the display screen was enlarged when I opened it back up.
thx for the file - karoke :)
The OCR window back-from-minimized should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
tormento
8th May 2020, 11:52
@Nikse555,
please of the feature list I posted, at least put Italic on the right side of the character to input during binary compare.
You don't know how many times I forget to enable/disable italic and delete part of the OCR database...
GCRaistlin
8th May 2020, 14:43
Bug:
Revoke write access to Waveforms directory for the current user.
Generate waveform data for a video file. You'll get an error.
Grant write access to Waveforms directory for the current user.
Press Retry.
Nothing happens. The 'Generate waveform data' window stays forever on the screen. No waveform data is being written to Waveforms directory.
GCRaistlin
8th May 2020, 17:48
Feature request: apply undo/redo to blocks of similar actions rather than to separate actions. For example, I adjust the boundaries of a subtitle using the waveform - by moving a boundary with the mouse. I perform one moving but SE considers it as a chain of small boundary shiftings. Hence, these actions replace the older ones in Undo stack. This is pointless as makes it harder to undo the whole action and makes it unable to undo the older actions.
I suggest to introduce a new setting: "Consider similar actions as separate if there are at least x seconds between them". Then, if x is set to 3, adjusts (e. g. by clicking on small arrows in 'Start time' or 'Duration' fields) if there were less than 3 secs between any two of them are considered as one action for Undo/Redo.
GCRaistlin
8th May 2020, 20:28
Feature request: keyboard shortcut for Video - Show/hide waveform.
GCRaistlin
8th May 2020, 20:57
Feature request: ability to set mouse wheel scroll step. Now it is 2 subtitles, I would like to have it set to 1 subtitle.
jlw_4049
9th May 2020, 04:37
thx for the file - karoke :)
The OCR window back-from-minimized should be fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Wanted to post back and say the issue with it zooming in to far on the 'subtitle image' window has not been solved with the latest Beta.
https://i.imgur.com/hX6gCfh.png
It is still doing this here when you minimize and come back. Although it seems like the program is reading them just fine.
EDIT: Maximizing the window seems to allow you to read it, however, staying in the smaller window it looks like the picture that I posted above.
Janusz
9th May 2020, 11:52
Wanted to post back and say the issue with it zooming in to far on the 'subtitle image' window has not been solved with the latest Beta.
@ Nikse555
I checked at home. The latest version 3.5.15 NEXT, beta 51 scales this window correctly.
For me, only the first stable version 3.5.15 had a problem with this.
It looked exactly like u jlw_4049.
tormento
9th May 2020, 12:37
Italian OCR correction wants to change
334
00:19:23,913 --> 00:19:25,081
ABBASSO I GLADIATORS
to
334
00:19:23,913 --> 00:19:25,081
ABBASSO i GLADIATORS
Nikse555
10th May 2020, 06:56
@tormento:
"ABBASSO I GLADIATORS" is not changed here... wrong language or something in your dictionaries?
Also, I'm not sure what you mean by "at least put Italic on the right side of the character to input during binary compare." - could you make a screenshot?
@jlw_4049/Janusz: I also cannot re-create the resize-and-restore-issue in latest beta, but I'll test on a few other computers.
jlw_4049, did you check version in Help -> About - also, how do you restore the minimized OCR window?
Latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Contains some good fixes for rippers:
- Bluray sup files could miss some images (where a subtitle would be expanded with more text)
- Teletext from .ts/.m2ts/.mts sometimes missed last subtitle
jlw_4049
10th May 2020, 07:04
@tormento:
"ABBASSO I GLADIATORS" is not changed here... wrong language or something in your dictionaries?
Also, I'm not sure what you mean by "at least put Italic on the right side of the character to input during binary compare." - could you make a screenshot?
@jlw_4049/Janusz: I also cannot re-create the resize-and-restore-issue in latest beta, but I'll test on a few other computers.
jlw_4049, did you check version in Help -> About - also, how do you restore the minimized OCR window?
Latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Contains some good fixes for rippers:
- Bluray sup files could miss some images (where a subtitle would be expanded with more text)
- Teletext from .ts/.m2ts/.mts sometimes missed last subtitleI'll try latest beta. Maybe I have an out dated version of the beta. Will double check in the AM.
It's been working perfectly other then that.
Will report back.
Sent from my Pixel 3a using Tapatalk
tormento
10th May 2020, 11:41
@tormento: "ABBASSO I GLADIATORS" is not changed here... wrong language or something in your dictionaries?
Same issue with:
871
01:00:00,263 --> 01:00:02,849
<i>Ripeto. I sospetti del Nite Owl
sono scappati.</i>
https://i1.lensdump.com/i/jBaLA7.md.png (https://lensdump.com/i/jBaLA7)
Here (https://www.mediafire.com/file/3oi17eue9oj1tfd/%5Bita%5D.7z/file) is the srt.
Fresh install. The OCR files are the ones you distribute.
Also, I'm not sure what you mean by "at least put Italic on the right side of the character to input during binary compare." - could you make a screenshot?
Here it is:
https://i.lensdump.com/i/jBaMVr.md.png (https://lensdump.com/i/jBaMVr)
tormento
12th May 2020, 18:08
It would be nice, when aborting OCR recognition, not to cancel the text of the current paragraph, but let it until the unrecognized character.
Sometimes it happens that some strange symbol can't be corrected by simply expanding and I have to abort to enter it manually. Unfortunately I have to enter the whole text!
Janusz
12th May 2020, 22:15
If we are already talking about it there is some inconsistency in the window operation
<Import/OCR Blu-ray (.sup)...> without consideration to the Selected OCR method.
Maybe someone so wanted so yes it works, but:
when the OCR process is stopped at the selected <Binary image compare>
or <OCR via nOCR> is as he wrote @tormento above.
When you select <Tesseract>, the line is recognized to the end of the
and only then the process is stopped.
the right side of the window and the 3rd list: <Unknow words>, <All fixes> and <Guesses used>.
When the OCR process works, these lists are populated accordingly.
When the process is stopped and resumed, the <Unknow words> list is cleaned completely,
and the other two do not. Therefore, always before the resumption of the process, I must first
check unknown words or correct errors in the <Unknow words> list before they disappear.
I think a better solution here would be to add to the list just as in the other two.
And ideally, in all 3 lists, the new text replaces the old from the line from which the process
was resumed and was not remarked at the end.
also in the window <VobsubOCRNOcrCharacter> not only in <VobSub - Manual image to text>,
wrote about it @GCRaistlin here. (https://forum.doom9.org/showpost.php?p=1909844&postcount=895)
The <Skip entire image> button could be useful, e.g. for illegible images and more.
Finally: There is an error in Polish translation to the program in line 2528:
is: <Skip>P&omoń</Skip>
to be: <Skip>P&omiń</Skip>
Melan
13th May 2020, 11:28
is: <Skip>P&omoń</Skip>
to be: <Skip>P&omiń</Skip>
And other error (line 2532):
<AutoSubmitOnFirstChar>Autom. proponuj &amp;pierwszy znak</AutoSubmitOnFirstChar>
<AutoSubmitOnFirstChar>Autom. proponuj pierwszy znak</AutoSubmitOnFirstChar>
borifax ;)
Nikse555
13th May 2020, 14:51
Same issue with:
[CODE]
https://i.lensdump.com/i/jBaMVr.md.png (https://lensdump.com/i/jBaMVr)
Good idea, fixed in latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Also, "Skip" in the OCR char window will now only skip from current character (and not the whole line).
@Melan/Janusz: thx - updated Polish translation.
(the "&" string will cause the following letter to be a shortcut - e.g. "&Skip" will react to the "Alt+S" shortcut).
Janusz
14th May 2020, 14:36
Note: Applies to version 3.5.15 NEXT, beta 92.
Thank you for this change, Nikse555.
Error creating <Unknow words> list.
https://drive.google.com/uc?export=view&id=1W5TpqxtLlct48xd2Fy9s7pQROFYuk2X1
Lines # 40, # 73 and # 81 - we have the word FBl there, and it has to be FBI.
I have already added the word FBI to the dictionary "names.xml" once.
By <Add pair to OCR replace list> I add FBl to FBI. I start OCR and I have it:
https://drive.google.com/uc?export=view&id=1mUqAarUdIUSzQIfZMeQ2RfJa8ZcrE4G1
In the <Subtitle text> window you can see that the conversion has been made and the word is known. This confirms the green color for this line.
Only that in <Unknown words> still hangs line # 40: FBl, although without # 73 and # 81.
Adding more word pairs works correctly - they do not appear again in the list. Well, unless there is no new word in the dictionary.
Line # 40 in this particular case will disappear only when I close the <Import / OCR Blu-ray ...> window and start the whole OCR process again.
But then another line with a different word will be the first forever with us until the window is closed.
I also checked it for words added to the dictionary - the first line displayed with the unknown word does not disappear.
Edition 1
The duplicate first lines will always appear on the second and subsequent file scans on all 3 lists also after changes made automatically
by the rules from the OCRFixReplaceList_User, OCRFixReplaceList files or with the option enabled <Fix common OCR errors ...] in Option/Settings/Tools.
They will not appear for automatic conversion of "l" (lowercase L) into "I" by Subtitle Edit, but we still don't see it on any of the lists,
except for an unknown word, when such a replacement creates a new incorrect word.
tormento
14th May 2020, 15:12
Also, "Skip" in the OCR char window will now only skip from current character (and not the whole line).
Thanks and please apply to abort too.
Nikse555
16th May 2020, 20:49
Latest beta has new (and hopefully improved) detection of space between italic letters: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Do let me know how it works! (it uses the value from "Set un-italic factor" in the list view context menu - probably normally between 0.22-0.32)
@tormento: thx for the test .sup files :)
Thanks and please apply to abort too.
I actually ment that it works for the "Abort" button ;)
@Janusz: I've fixed an issue related to your last post, but it's really hard to test without your exact setup/sup... could you make a .zip archive with all relevant files, if latest beta still has issues?
GCRaistlin
16th May 2020, 23:22
When performing OCR it is unable to add a proper match for the percent sign (https://mir.cr/10PHMJUD, # 68): SE recognizes its first part as "o". To add a better match, I deleted this "o" from the DB and run OCR again. This time the first part was recognized as "O", and 'Delete' button is inactive.
jlw_4049
17th May 2020, 00:52
The latest BETA struggles with ♪ characters very badly.
Nikse555
17th May 2020, 06:41
When performing OCR it is unable to add a proper match for the percent sign (https://mir.cr/10PHMJUD, # 68): SE recognizes its first part as "o". To add a better match, I deleted this "o" from the DB and run OCR again. This time the first part was recognized as "O", and 'Delete' button is inactive.
thx for the file :)
To fix "%" double click in the list view in main OCR window, then right-click in the list box in the "Inspect" windows and choose "Add better multi match", then expand the images to cover the "%" sign:
https://nikse.dk/ocr-percent.png
The latest BETA struggles with ♪ characters very badly.
I probably need more info... subtitle + screenshots... you're using Tesseract for OCR'ing?
Nikse555
17th May 2020, 08:25
Shortcuts for the "OCR Character" window is:
https://nikse.dk/se-ocr-char.png
Expand selection: Alt + arrow right
Shrink selection: Alt + arrow left
Toggle italic: Ctrl+I (+ Alt+I depending on translation)
Toggle auto-submit-first-char: Alt+F (depending on translation)
Skip current letter(s): Esc (+ Alt+S depending on translation)
Skip entire subtitle: Ctrl+Shift+S (new shortcut)
tormento
17th May 2020, 11:37
Shortcuts for the "OCR Character" window is
Now that you made me think about it, it would be nice to have the capability to expand both right and/or left side. Sometimes % character has bad OCR on the left and/or on the right too. The only think I can do now is abort and input it manually. Two buttons such as
|←|expand|→|
|→|shrink|←|
or the same changing function with SHIFT key would be nice.
P.S: The red italic word on the right of the character is great. You could remove the one on top of the window now. :)
GCRaistlin
17th May 2020, 22:39
What does auto-submit-first-char do?
Nikse555
17th May 2020, 23:02
What does auto-submit-first-char do?
It will use the first key down as the OCR letter without waiting for a click on "OK" or the "Enter" key pressed.
I often set the error rate to zero when starting OCR of a new sub for the first 10-20 lines, in which case I add a lot of single letters, and that's much faster without having to press the "Enter" key or the "OK" button.
(you need to turn it off again, if the prompt is for a multi letter image, like "ft")
GCRaistlin
17th May 2020, 23:11
Nikse555
Thanks. I'd say it is needed to add a brief explanation for this option to the UI, as long as for "add better multi match", as these options' names aren't self-explanatory.
GCRaistlin
17th May 2020, 23:25
thx for the file :)
To fix "%" double click in the list view in main OCR window, then right-click in the list box in the "Inspect" windows and choose "Add better multi match", then expand the images to cover the "%" sign:
Something went wrong. I performed all the actions above, then rerun OCR from this line - SE didn't ask me anything but the percent sign is missing in the recognized line:
https://i112.fastpic.ru/thumb/2020/0518/f0/f29c6ed49b180fc586df360f20e75cf0.jpeg (https://fastpic.ru/view/112/2020/0518/f29c6ed49b180fc586df360f20e75cf0.jpg.html)
UPD: It seems that I didn't enter "%" to the field. It's worth to check if it isn't empty...
GCRaistlin
17th May 2020, 23:47
Bug(s):
Follow the steps above but add a wrong match, for example "@".
Start OCR from the same line, then interrupt it.
Call 'Inspect compare matches' window.
Delete the wrong match, add the right match, press OK.
Start OCR from the same line again.
You'll get 'VobSub - Manual image to text' window for the char you have just added a match for. And by the way the window title is incorrect - it's not the VobSub being recognized. But let's go further.
Press Abort, try to add multi match again. You'll get 'Image already in db' error.
Janusz
18th May 2020, 03:04
@Janusz: I've fixed an issue related to your last post, but it's really hard to test without your exact setup/sup... could you make a .zip archive with all relevant files, if latest beta still has issues?
I use Windows 10, 64 bit. For this test Subtitle Edit 3.5.15 NEXT, beta 106, and nOCR.
For the purposes of the test I am not using pol_OCRFixReplaceList_User.xml.
Settings.xml, pol_OCRFixReplaceList, test.sup, test.db (incomplete), test.nocr in janusz.test.zip for download. (https://drive.google.com/uc?export=view&id=1RkRsnLSHtUeOoArF9KH9Y9T7Bd9ioX5e)
Images for this file come from various subtitles, hence many duplicate characters, but this is not a problem.
These or other characters are to be interpreted (read) correctly. This is the assumption.
I know that the sup file for this test was generated from images containing some error and the number of errors
(5 in 4 lines) has nothing to do with the number of errors in the text consisting of a thousand or more lines.
https://drive.google.com/uc?export=view&id=1usEFhkMjzJz_1TuzgyGO7ioiY0-DxXn7
To begin with, the analysis of text created without a dictionary - that is, how the program itself deals with OCR.
1. Lines 5, 6 and 7 we see "I", which we do not have in the character database. Creating the base for this example
was not possible because the text does not contain "I" at all. So where does this come from? Suspicion falls on the program.
Browsing this forum, not everyone looks here, we'll find out that the program can replace l with I: at the beginning,
in the middle and, surprisingly, at the end of words written in lower case.
And also at the beginning of a paragraph or task - example line 5 where the dot in this case does not mean the end of the sentence.
Lines 6 and 7 in the original texts were a continuation of the sentence and should not be changed.
I will add that the words "lub" (or), "lecz" (but) are used in Polish often so for normal, full text there will be many mistakes.
I believe that the program function, which always works, cannot generate errors for any selected language, especially in its absence.
What have we gained? 3 errors instead of 0 (zero). With longer texts, the number of good replacements will always be less than
the number of errors for a simple reason. Statistical "I" is less common than "l" at the beginning of words, and certainly not
in the middle or end of words written in lowercase. Therefore, I would prefer to correct only errors arising in the OCR process.
Why do I need extra?
2. Line 8. There are 2 cases of combined words here. "chybajuż" and "przynajmniejpod".
I can improve them by reducing [No of pixels is space] to 3. I will get "przynajmniej pod" - that's OK, the rest of the text above.
The phrase "chybajuż" will divide into two words "chyba już" only at 2. However, now OCR found additional apostrophes,
which at the beginning creating a character base I combined into one ["].
The effect: line 8 is OK, but the text above went apart. [No of pixels is space] parameter is too small,
hence my request in one of the previous posts for a different space for italics.
Interesting fact: selecting [Inspect nocr matches for current image ...] on line 8 will display the text correctly with appropriate spacing for [No of pixels is space] = 2, pressing OK will not save any changes to the text, however, because this window is only for characters in the database. If selecting OK saved these changes to the text would be great, at least until the italics problem is solved globally.
OKAY. To deal with line 8 I return to setting [4]. I switch the dictionary to Polish.
The pol_OCRFixReplaceList.xml file already contains a <WordPart from = "j" to = " j" /> line in the <PartialWords> section
- this is OK for "chybajuż" - but let's see what happened with "przynajmniejpod".
Based on a comment to this section: the program added a space before "j", did not find in the dictionary either "przynajmnie"
or "jpod" - such words do not exist in Polish. For me, the repair program should end its work at this stage and change nothing.
Why did he divide the program by "j" and also replace "p" with "j". I could use <WordPart from = "j" to = "j "> for this and similar expressions,
but such a conversion in at least the Polish language will divide one correct word into two other also correct, e.g. "najjaśniejszy" (brightest)
to "naj" (most) and "jaśniejszy" (brighter) so I can't use it. In addition, I will not see such a replacement on the [All fixes] list or on [Guesses used] as opposed to substituting "p" for "j". This replacement is visible and can be quickly corrected manually.
Bottom line: it remains to improve "improved" again, as in item 1.
I saw in some files, e.g. dan_OCRFixReplaceList.xml, in the part concerning division into two words, such a notation,
e.g. <WordPart from = "o" to = "e" />. Why is this supposed to serve as not just a simple conversion of "o" to "e".
3. Now lines 1 to 4. They look flawless - that's how it is. Please perform [New], we will create a new character base,
any name other than "test", press [Edit], [Import] - indicate our base "test.nocr", [OK].
We return to OCR, we set ourselves on the first line and [Start].
Result: during import we lost all characters resulting from the combination of 2 or 3 adjacent characters, i.e. [''] is ["], [o/o] is [%].
This is what it looks like. Characters added by the extension to the adjacent character or characters,
are invisible in the database once, and two are lost when importing into a new character base.
It's a lot, but I wanted to write more than just "not working".
Thank you for the dark background in [Set un-italic factor].
I wanted to ask for this for a long time.
varekai
18th May 2020, 12:52
Something went wrong. I performed all the actions above, then rerun OCR from this line - SE didn't ask me anything but the percent sign is missing in the recognized line:
https://i112.fastpic.ru/thumb/2020/0518/f0/f29c6ed49b180fc586df360f20e75cf0.jpeg (https://fastpic.ru/view/112/2020/0518/f29c6ed49b180fc586df360f20e75cf0.jpg.html)
UPD: It seems that I didn't enter "%" to the field. It's worth to check if it isn't empty...
WTF!! You are extremely obnoxious!
Are you really that stupid?
If you wanna post a link to your neverending images use imgur.com and point directly to the jpg
https://i.imgur.com/7jpu1si.jpg
or use imgur.html
https://imgur.com/a/dOBSfAA
Please stop using fastpic*ru it's awful!!
Grr...
Nikse555
18th May 2020, 14:50
@Janusz: Sorry, I've not really done any work with "nOcr (line ocr)"... I've mostly done stuff to improve "Binary image compare"
With latest beta ( https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip ) I get this result with your sup file:
https://nikse.dk/se-doom9-pol.png
Janusz
18th May 2020, 15:20
Thank you, Nikse555.
I also thought that nothing was happening with nOCR.
Please, look again at line 8 at home and my attention 2 above. Why this division and why is "p" converted to "j"?
The nOCR method gave me the same result with latest beta 119
Nikse555
18th May 2020, 15:48
@Janusz: yes, thx :)
line 8 seems to be a bug - I'll look into it.
Nikse555
18th May 2020, 16:35
@Janusz: Beta updated: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
(also tried to fix the nOcr issues)
Works best with pixes-is-space = 3 for me...
Bug(s):
Follow the steps above but add a wrong match, for example "@".
Start OCR from the same line, then interrupt it.
Call 'Inspect compare matches' window.
Delete the wrong match, add the right match, press OK.
Start OCR from the same line again.
You'll get 'VobSub - Manual image to text' window for the char you have just added a match for. And by the way the window title is incorrect - it's not the VobSub being recognized. But let's go further.
Press Abort, try to add multi match again. You'll get 'Image already in db' error.
Thx, should also be fixed in above beta.
jlw_4049
18th May 2020, 16:36
I will try next beta out [emoji846]
Sent from my Pixel 3a using Tapatalk
Nikse555
18th May 2020, 18:50
@Janusz: And now really fixed the italic-space-stuff in nOcr: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Janusz
18th May 2020, 21:07
@Janusz: And now really fixed the italic-space-stuff in nOcr: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
It's perfect now. Two lines in pol_OCRFixReplaceList.xml
<WordPart from = "ą" to = "ą " />
<WordPart from = "j" to = " j" />
divide expressions consisting of two or even three combined words into single words. With an earlier amendment regarding "l" and "I", the text consisting of 1189 lines, of which almost half was written in italics, is read almost 100%. There are two mistakes to improve. If you add them to your replacements, the effectiveness will be 100%.
Really good work @ Nikse555. Thank you again.
GCRaistlin
18th May 2020, 22:08
Nikse555
The latest beta still allows to add an empty better multi match.
Could you please allow selecting a character by a right click in 'Inspect items' area of 'Inspect compare matches for current image' window? I mean along with showing the context menu.
varekai
19th May 2020, 07:05
This is the message that was sent from you:
***************
Who are you to speak to me this way?
These pics aren't for you.
Don't open them and relax if you never have heard about ad blockers.
***************
This is an open forum, when you post something, everyone can read/see your post.
We post images, to show and clearify the issues we have and want to report bugs, get input/help from the forum.
You make a clickable link and the whole idea of that is to make someone click on it, right?
Therefore you should spare the forum from ugly sites like fastpic*ru.
Of course I have many layers of protection for my computer, including AV, Firewall, AD- and Script-blockers and what not.
Not all in here have that protection.
How hard can it be for you to understand that? Really?
Do yourself and the forum a favour and use another imagehost, imgur*com is very good and easy to use and... it's adfree (almost)!
No nasty pics close to pron, no ads, no popup windows etc etc...
If you don't understand the difference...
This is it:
Link:
https://imgur.com/a/Lqb3rjn
Image:
https://i.imgur.com/DPDnOaG.png
varekai
19th May 2020, 11:33
@GCRaistlin
This is the message that was sent from you:
***************
If you don't understand that other visitors aren't interested in this discussion it's your problem.
Nobody else seems to care about Fastpic so don't try protecting those who don't need your protection.
And don't bother to address me again on the forum, you won't get any answer.
***************
tormento
19th May 2020, 11:54
Latest beta has new (and hopefully improved) detection of space between italic letters:
Enjoy with line 13 of this (https://www.mediafire.com/file/2lkuf2xhlxkp3vv/Apollo_13_eng.7z/file). :)
Melan
19th May 2020, 12:39
https://i.imgur.com/oePIeyR.png
I did it in 10 minutes.
http://www.mediafire.com/file/7daaf4nb889ffak/Apollo_13_eng.srt/file
tormento
19th May 2020, 16:45
I did it in 10 minutes.
And you did it wrong. :D
"of" is italic while in your OCR it is in normal style.
I am finding issues, not establishing OCR time records.
Would you please explain me how can line 1097 contain the {\an8} marker?
I never noticed Subtitle Edit was capable of it.
Janusz
19th May 2020, 20:38
"of" is italic while in your OCR it is in normal style.
To make the text look good, [No of pixels is space] = 12, and this means that "of Apollo" is one word "ofApollo" and as such it was probably included by the algorithm as not italics. I think so - I don't know the algorithm. I do not know at what moment it is divided into two words, or on what terms. Probably this happens after selecting the "English" dictionary and selecting: [Fix OCR errors] and [Try to guess unknown words].
If you change [No of pixels is space] to e.g. 8, you will get 2 words "of" - in italics and "Apollo" - not italics, and "</i>" will be inserted after "of", but with such a small space remaining text will split up.
As you can see, this functionality still needs to be refined.
Would you please explain me how can can 1097 contain the {\ an8} marker?
I never noticed Subtitle Edit was capable of it.
For some time this Subtitle Edit tag added to me while importing subtitles from ts files for texts placed at the top of the screen. I don't remember which version.
Perhaps at this time other permanently embedded subtitles will appear at the bottom of the screen.
Melan
19th May 2020, 20:51
@tormento
Don't be a child. If you think that more than 2,000 lines will not contain errors, you are wrong.
SE works really well.
Nikse555
20th May 2020, 12:20
@Melan: thx, I think SE works really well too. It's still nice with feedback and ideas as it might help with making SE even better.
@tormento: Ah, did you set the proper "italic factor"? Right click in the list view, and choose "Set un-italic" factor (I think it's called). [No of pixels is space] = 13 worked fine for me I think.
SE can detect top align from Bluray .sup files - can be toggled via right click on the image... I've also added a on-video-preview for each image - press Ctrl+P to see the subtitle on actual screen size.
@GCRaistli
>The latest beta still allows to add an empty better multi match.
I think that "empty string" could be a valid text... perhaps a warning?
>Could you please allow selecting a character by a right click in 'Inspect items' area of 'Inspect compare matches for current image' window?
I don't follow... ?
Latest beta: https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
Janusz
20th May 2020, 14:02
@Nikse555
Is there sense for the nOCR method to continue reporting bugs in this forum since no one is using this method here?
As you wrote above, you recommend "Binary image compare", and nothing has happened with the nOCR project for a long time.
I would just ask you to fix the crash of the nOCR process from the start when "no dictionary" was selected.
I get this error (last beta 123 and several earlier) regardless of the configuration for the program.
In stable versions 3.5.14 and 3.5.15 this error is not there. If you need any files, you can use those from 18/05/2020.
Setting various options except [Dictionary = none] in the nOCR window does not affect the error.
https://drive.google.com/uc?export=view&id=1oPsaPYKK8HOOhclbrik2UIdkof-Sr0QG
Excerpt from error_log.txt
----------------------------------------------- ------------------------------
Date: 05/19/2020 22:38:29
Message: Unable to load '' (also check libc.so.6 + libdl.so.2)
-------------------------------------------------- ---------------------------
Date: 05/19/2020 22:38:29
Message: Not all required methods was found in libvlc
-------------------------------------------------- ---------------------------
Date: 05/19/2020 22:52:21
Message: Unable to load '' (also check libc.so.6 + libdl.so.2)
-------------------------------------------------- ---------------------------
Nikse555
20th May 2020, 14:27
@Janusz: Is the crash fixed in this beta?
https://github.com/SubtitleEdit/subtitleedit/releases/download/3.5.15/SubtitleEditBeta.zip
vBulletin® v3.8.11, Copyright ©2000-2026, vBulletin Solutions Inc.