Superb quality and spec AB-Com PULSe 4K SE. Crazy offer! Only £99! FREE UK DELIVERY! 4K UHD, Enigma 2, Multiboot 4 images & more!...
Superb quality and spec AB-Com PULSe 4K Rev II Twin Satellite tuner only £149! FREE UK DELIVERY! 4K UHD, Enigma 2, SATA HDD facility, Multiboot 4 images & more!...

[ViX_Misc] Autotimers and Description Uniqueness

Yes AFAICS in a running image we have /usr/lib/enigma2/python/Plugins/Extensions/AutoTimer/AutoTimer.pyo which we can rename and replace with a modified AutoTimer.py, new new AutoTimer.pyo will be created the first time it's used and then you could delete the AutoTimer.py if you wished.
You need to restart the GUI to get the pyo file rebuilt. If there are any errors, the pyo file won't be there.
 
Yes, there is something more than I described going on. Most likely due to bugs/errors elsewhere I'm thinking?

I don't think maximum matching sequence length ratio is right for this at all, even if you somehow manage to shuffle things around first
I just haven't got my head around exactly what would be better yet.

(Oh by the way it is configured to ignore spaces, I think)

After looking at the code there are two sepearate things happening with these problem EPGs

There are 3 separate tests working on 3 discrete bits of EPG.
I) the title data
ii) the short description data
iii) the extended description data

Test 1: only the title data is compared

Test 2: only the short description data is compared,but only if:
i) test 1 produced a match
ii) the menu option “title and short description” has been selected

Test 3: only the extended description data is compared but only if:
i) test 2 produced a match
ii) the menu option “title and all descriptions” has been selected

Test 2 is falling over on these problem EPG beacuse it is identifying all programs with a genric description with only the last few charcters changing to be similar enough to be identical.

Test 3 is falling over because there is no extended description data to check but when checking this non-existant data it indicates that every time there is a difference. Garbage in = garbage out - or more correctly garbage in = the same result out every time. This is overriding the result of test 2. Note: there may be extended description from the broadcasters in other countries and maybe if the epg information is obtained over the net.

If it can be confirmed that there is no extended description in the UK then UK users possibly should not select the "all descriptions" in the menu beacuse even if test 2 was changed to detect very small differences test 3 may overwrite the results with garbage with respect to setting a timer.

Currently this still leaves the original problem in that the only way of getting every episode with no repaets for programs with generic EPGs is knowing what time they are being broadcast and setting filters to limit the time slot.
 
Last edited:
Require description to be unique: Any service/recording
Check for uniqueness in: Tile and all descriptions
I've definitely had cases where this works.
I think maybe, for me, it sometimes works for Sky, but not for Freeview.
 
I installed a modified AutoTimer.py with print statements to log what's going on a bit when I run an AutoTimer scan.
I'm attaching the AutoTimer.py and the log here (in a ZIP file because the log is BIG) if anyone wants a look, or to improve on what I've done.
As of yet I haven't really made much sense of it since I didn't set any particular timers.
I'll just note here that "The Simpsons" is an AutoTimer with my, now infamous, settings looking only on Channel 4 HD on Sky.
And "Malcolm in the Middle" is an AutoTimer with my much maligned settings looking only on 4Music on Freeview.
I have to go and eat now.
 

Attachments

I installed a modified AutoTimer.py with print statements to log what's going on a bit when I run an AutoTimer scan.

I added a few more debug print lines and set an autotimer (uniqueness in title and all descriptions) and got inconsistant results.

i) The short description always contained the whole of the description including the [s1, Ep3] bit at the end of a common generic description but short description test allways found them to be identical despite having different series and episode numbers

Sometimes the extended description conatained an exact copy of the short description and sometime it was blank. With a duplicate description extended test gave the same answer as the short description test,, With a blank externded description the test always assumed they were different.

So if the descriptions only vary by a few charaters they will be assumed to be identical.
If the all descriptions filter option is chosen and the extended data is poulated the test will return the same results as the short description test. If the extended data is not populated the test will always return a false result.

Judging by what I've now seen its a lottery if the extended description data is populated. Check for uniqueness in Title and Short decription will probably work as intended for 99% of programs but all uniqueness options will fall down when the description for all episodes and reapeats is virtualy the same apart from a few charactes.

3 samples showing what is in the title data, short description data and extended description data

---BEGIN---
AutoTimer name1 ='Two in Clover'
AutoTimer name2 ='Two in Clover'
AutoTimer shortdesc1='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep2/6.'
AutoTimer shortdesc2='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer extdesc1 ='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep2/6.'
AutoTimer extdesc2 ='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer foundTitle ='True'
AutoTimer foundShort ='True'
AutoTimer retValue ='True'
---END-----

---BEGIN---
AutoTimer name1 ='Two in Clover'
AutoTimer name2 ='Two in Clover'
AutoTimer shortdesc1='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer shortdesc2='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer extdesc1 ='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer extdesc2 ='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep1/6.'
AutoTimer foundTitle ='True'
AutoTimer foundShort ='True'
AutoTimer retValue ='True'
---END-----

---BEGIN---
AutoTimer name1 ='Two in Clover'
AutoTimer name2 ='Two in Clover'
AutoTimer shortdesc1='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep3/6.'
AutoTimer shortdesc2='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep4/6.'
AutoTimer extdesc1 =''
AutoTimer extdesc2 =''
AutoTimer foundTitle ='True'
AutoTimer foundShort ='True'
AutoTimer retValue ='False'
---END-----
 
i) The short description always contained the whole of the description including the [s1, Ep3] bit at the end of a common generic description but short description test allways found them to be identical despite having different series and episode numbers
No, it didn't find them to be identical. It found them to be similar.

The relevant code is in checkSimilarity() (note the hint in the name).

With a blank externded description the test always assumed they were different.
That is how it's coded, although that looks like a bug.
If you are looking for similarities in Long Descriptions (so searchForDuplicateDescription == 2) it then checks for extdesc1 and extdesc2. If these are unset (or empty strings) that is False. So it never runs a similarity check on them and retValue is left as False, which gets returned.
I suspect an else: retValue = True is needed....
 
!!!!OOPS PASTED WRONG THING IN IGNORE UNTIL I FINISH EDITING IT!!!!!



---BEGIN---
AutoTimer name1 ='Two in Clover'
AutoTimer name2 ='Two in Clover'
AutoTimer shortdesc1='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep3/6.'
AutoTimer shortdesc2='1970s comedy series with Sid James and Victor Spinetti as two city-dwellers who have moved to the countryside - expecting farming life to be simpler. S2, ep4/6.'
AutoTimer extdesc1 =''
AutoTimer extdesc2 =''
AutoTimer foundTitle ='True'
AutoTimer foundShort ='True'
AutoTimer retValue ='False'
---END-----
Ah. I think I see it.
In the original AutoTimer.py if we get as far as comparing extdesc1 and extdesc2 there is an if statement testing if either one is an empty string, we should then always return True.
But the else part is missing, there's another else underneath but I don't think that applies, it certainly wouldn't in other programming languages I've been taught.
I'm going to try this:
Code:
	def checkSimilarity(self, timer, name1, name2, shortdesc1, shortdesc2, extdesc1, extdesc2, force=False):
		foundTitle = False
		foundShort = False
		retValue = False
		if name1 and name2:
			foundTitle = ( 0.8 < SequenceMatcher(lambda x: x == " ",name1, name2).ratio() )
		# NOTE: only check extended & short if tile is a partial match
		if foundTitle:
			if timer.searchForDuplicateDescription > 0 or force:
				if shortdesc1 and shortdesc2:
					# If the similarity percent is higher then 0.7 it is a very close match
					foundShort = ( 0.7 < SequenceMatcher(lambda x: x == " ",shortdesc1, shortdesc2).ratio() )
					if foundShort:
						if timer.searchForDuplicateDescription == 2:
							if extdesc1 and extdesc2:
								# Some channels indicate replays in the extended descriptions
								# If the similarity percent is higher then 0.7 it is a very close match
								retValue = ( 0.7 < SequenceMatcher(lambda x: x == " ",extdesc1, extdesc2).ratio() )
							else:				# Brian was here
								retValue = True		# Brian was here
						else:
							retValue = True
			else:
				retValue = True
		return retValue

Now perhaps all we have to deal with is the way just one character change close to the end of part of the description is ignored?

I'm currently thinking that the Levenshtein ratio makes more sense than the SequenceMatcher.ratio function currently used.
But so far I can't find an implementation that doesn't use things I don't seem to be allowed to import.
Also I think Numbers should be make more important than other characters so that changes in numbers are not ignored.
 
Last edited:
Drat, and other words I'm not allowed to use.
Ignore the last post and try and excuse me for being an idiot.
I'm giving up for tonight.

By the way I'm currently thinking that the Levenshtein ratio makes more sense than the SequenceMatcher.ratio function currently used.
But so far I can't find an implementation that doesn't use things I don't seem to be allowed to import.
Also I think Numbers should be made more important than other characters so that changes in numbers are not ignored.
 
By the way I'm currently thinking that the Levenshtein ratio makes more sense than the SequenceMatcher.ratio function currently used.
You do seem to be missing the real point.
It's very simple for a human to read two similar texts and decide whether they are, in fact, the same - we've evolved over thousands of years to recognize patterns.
There is no way that simple computer code can achieve the same discrimination.
No matter what comparison code you put in place it will never come up with the "correct result" every time.
So live with that limitation and look for some some other test to help you.
 
You do seem to be missing the real point.
It's very simple for a human to read two similar texts and decide whether they are, in fact, the same - we've evolved over thousands of years to recognize patterns.
There is no way that simple computer code can achieve the same discrimination.
No matter what comparison code you put in place it will never come up with the "correct result" every time.
So live with that limitation and look for some some other test to help you.

That all makes sense Birdman but why is it Sky boxes or Humax boxes etc never get it wrong?
 
That all makes sense Birdman but why is it Sky boxes or Humax boxes etc never get it wrong?

Maybe because they use series CRID data.

Some progress on enigma2/CRID's achieved in Australia a while ago ....

Code:
http://beyonwiz.com.au/forum/viewtopic.php?f=54&t=10773
 
Last edited:
.... this link looks quite interesting....

Code:
https://www.beyonwiz.com.au/forum/viewtopic.php?p=183023#p170895

@IanSav is probably the best bet for comments.
 
You do seem to be missing the real point.
It's very simple for a human to read two similar texts and decide whether they are, in fact, the same - we've evolved over thousands of years to recognize patterns.
There is no way that simple computer code can achieve the same discrimination.
No matter what comparison code you put in place it will never come up with the "correct result" every time.
So live with that limitation and look for some some other test to help you.
Oh come on. I never said my answer would be perfect.

Sent from my SM-A515F using Tapatalk
 
I contributed to a few threads discussing CRIDs several years ago and actually raised a request here on Vix to see if CRID handling could be incorporated into the AutoTimer.py code but that request thread went nowhere at the time. The beyonwiz team in Australia (as per @ccs post) have had some success in using this CRID data, but I think the big issue is the lack of consistency of how the CRID information is structured or encoded. (as @ccs points out, IanSav (and prl) have worked on this in the Australian environment). What might work for Freesat/Freeview might need workarounds or kludges for SKY or other broadcasters. I used the "Series Link" feature on SKY and on Humax Freesat receivers and found them excellent. I originally thought the AutoTimer mechanism on enigma to be a little crude in its operation but, over the years I've adapted to its foibles and have found that it works for me 99% of the time. I rarely miss episodes I want to record and it generally works when a new series starts after being off the air for months (providing the broadcaster doesn't move it to a completely new channel or time slot). Worst case scenario is that I have multiple recording of the same episode occasionally.
The CRID mechanism would eliminate the ambiguity of trying to match on title or description but I imagine it would need a complete overhaul of the program logic to keep tables of series and programme CRID info to determine if particular episodes need to be recorded or have already been recorded.
 
Last edited:
... I'm sure there would be room in the *.ts.meta files to store an extra word or two of crid details.

If it's blank/missing, use the existing system, if it's not, bingo.
 
I contributed to a few threads discussing CRIDs several years ago and actually raised a request here on Vix to see if CRID handling could be incorporated into the AutoTimer.py code but that request thread went nowhere at the time. The beyonwiz team in Australia (as per @ccs post) have had some success in using this CRID data, but I think the big issue is the lack of consistency of how the CRID information is structured or encoded. (as @ccs points out, IanSav (and prl) have worked on this in the Australian environment). What might work for Freesat/Freeview might need workarounds or kludges for SKY or other broadcasters.

This part of the problem in that any changes have to work for all broadcasters irrespective where in the world they may be. Even in the UK the main channels may have a good record with CRID data but there have in the past also many instances on the "lesser" channels where strict transmitting of the correct CRID has been a bit lax.

I originally thought the AutoTimer mechanism on enigma to be a little crude in its operation but, over the years I've adapted to its foibles and have found that it works for me 99% of the time. I rarely miss episodes I want to record and it generally works when a new series starts after being off the air for months (providing the broadcaster doesn't move it to a completely new channel or time slot).

I've also found Autotimers to be reliable 99+% of the time and if anything "goes wrong" it tends to record too much rather than missing recordings. I can live with the occsaional 2 or 3 copies of the repeat. I tend not to set limited time slots so when a program does move time it tends to be captured three months down-line.

Note: all my autotimers are set to check in the title and short description only.
 
The comparisons of the titles and descriptions seem to be done in function checkSimilarity which starts on line 838 of the file AutoTimer.py.
It uses a function SequenceMatcher from difflib, which you can find descriptions of on the web such as https://towardsdatascience.com/sequencematcher-in-python-6b1e6f3915fc

I don't think it's really the right function in this application because, for instance, a single character difference in the middle of a description counts as a huge difference while a single character difference near the beginning or end counts only as a small difference. Thus the change from (S01:E04) to (S01:E05) right at the end of a description is seen as something to ignore.

Okay I've done some tests and THIS IS WRONG.
It does not seem to see differences in the middle as more important than differences near the beginning and end.
The descriptions of the SequenceMatcher function I found seem over simple and use over simple examples so I didn't understand exactly what it does (and I still don't).
SORRY.:o

SequenceMatcher probably is a good choice except that numbers need to be given more importance, and I have an idea for that which I will try soon.
Other people are, as always, free to ignore what I write.
 
SequenceMatcher probably is a good choice except that numbers need to be given more importance, and I have an idea for that which I will try soon.
Other people are, as always, free to ignore what I write.


But don't forget there may be many numbers in the description that don't relate to the series or episode. I saw one description the other day it said something like "....post war britain between 1943 and 1952........" Perhaps just making numbers more important in the current tests is not the way to go.

Also don't forget that any solution for one problem cannot create another and break something the autotimer does well. Fixing a problem in less than 1% of descriptions cannot cause prolems with, say, 3% of other descriptions.
 
But don't forget there may be many numbers in the description that don't relate to the series or episode. I saw one description the other day it said something like "....post war britain between 1943 and 1952........" Perhaps just making numbers more important in the current tests is not the way to go.

Also don't forget that any solution for one problem cannot create another and break something the autotimer does well. Fixing a problem in less than 1% of descriptions cannot cause prolems with, say, 3% of other descriptions.

Okay, lets confine ourselves to fixing the big logical error you described:
There are 3 separate tests working on 3 discrete bits of EPG.
I) the title data
ii) the short description data
iii) the extended description data

Test 1: only the title data is compared

Test 2: only the short description data is compared,but only if:
i) test 1 produced a match
ii) the menu option “title and short description” has been selected

Test 3: only the extended description data is compared but only if:
i) test 2 produced a match
ii) the menu option “title and all descriptions” has been selected

Test 2 is falling over on these problem EPG beacuse it is identifying all programs with a genric description with only the last few charcters changing to be similar enough to be identical.

Test 3 is falling over because there is no extended description data to check but when checking this non-existant data it indicates that every time there is a difference. Garbage in = garbage out - or more correctly garbage in = the same result out every time. This is overriding the result of test 2. Note: there may be extended description from the broadcasters in other countries and maybe if the epg information is obtained over the net.
The code is clearly not supposed to do this, there is a test for it, but it's screwed up. Maybe the code I posted before and then lost confidence in is the fix:
Code:
	def checkSimilarity(self, timer, name1, name2, shortdesc1, shortdesc2, extdesc1, extdesc2, force=False):
		foundTitle = False
		foundShort = False
		retValue = False
		if name1 and name2:
			foundTitle = ( 0.8 < SequenceMatcher(lambda x: x == " ",name1, name2).ratio() )
		# NOTE: only check extended & short if tile is a partial match
		if foundTitle:
			if timer.searchForDuplicateDescription > 0 or force:
				if shortdesc1 and shortdesc2:
					# If the similarity percent is higher then 0.7 it is a very close match
					foundShort = ( 0.7 < SequenceMatcher(lambda x: x == " ",shortdesc1, shortdesc2).ratio() )
					if foundShort:
						if timer.searchForDuplicateDescription == 2:
							if extdesc1 and extdesc2:
								# Some channels indicate replays in the extended descriptions
								# If the similarity percent is higher then 0.7 it is a very close match
								retValue = ( 0.7 < SequenceMatcher(lambda x: x == " ",extdesc1, extdesc2).ratio() )
							else:			# Brian was here
								retValue = True	# Brian was here
						else:
							retValue = True
			else:
				retValue = True
		return retValue
 
... I'm sure there would be room in the *.ts.meta files to store an extra word or two of crid details
Not where it's needed. You really want to know you've already recorded something even after you've deleted the timer for and the recording of it.
A sqlite database of all recorded CRIDs would be the thing to use.
With a configurable "forget after" time, so that any record older then this would be pruned.
 

OpenViX Feeds Status

Back
Top