Superb quality and spec AB-Com PULSe 4K SE. Crazy offer! Only £99! FREE UK DELIVERY! 4K UHD, Enigma 2, Multiboot 4 images & more!...
Superb quality and spec AB-Com PULSe 4K Rev II Twin Satellite tuner only £149! FREE UK DELIVERY! 4K UHD, Enigma 2, SATA HDD facility, Multiboot 4 images & more!...

[OS-mio] Crash when folder contains filenames with Windows encoding

Avoiding the crash is simple, for example modify getItemDisplayName() to return value: name.encode('UTF-8', 'surrogateescape').decode('UTF-8', 'ignore')
Why would the encode be needed? Isn't name starting as a bytes item (given that this is the only data-type which can contain all valid filenames)?

Now question is what name / characters it should display to user? But I think display problem is only minor issue here.
What about:

Code:
 name.decode(encoding='utf8', errors='backslashreplace')
which at least shows what was there.
 
Assuming we are still looking at filenames in some language encoding then…..
We have no idea what "language" encoding may be used in a filename. A filename is just a series of bytes - any bytes. A file may have come from anywhere.
 
So maybe... encode (with surrogates) to get a bytes string... then try to get the encoding with chardet... then use that encoding in the decode.
Yes, works for displayname. Maybe because data.txt is inside structure and not a function parameter.

For example this works for me:

s = 'p\udcf6ll\udcf6.mp4' #str
b = s.encode('UTF-8', 'surrogateescape') #bytes
import chardet
e = chardet.detect(b)['encoding']
s2 = b.decode(e)
print(s2) # Displays correctly for me

Btw, just noticed "getItemDisplayName" is imported also in MovieSelection.py and it uses it for file operations?? Function name "DisplayName" suggest it should be used only for the display.

It would be better to have another function "getItemFileName" and that return bytes, name.encode('UTF-8', 'surrogateescape') and adjust code in movieselection.
 
Why would the encode be needed? Isn't name starting as a bytes item (given that this is the only data-type which can contain all valid filenames)?

Currently function getItemDisplayName return value "name" is STR, not bytes. You can convert it to bytes: name.encode('UTF-8', 'surrogateescape'). But then every code that uses "getItemDisplayName" should be adjusted for bytes. It is possible to do this also. Question is only what is the best way
 
So maybe... encode (with surrogates) to get a bytes string... then try to get the encoding with chardet... then use that encoding in the decode.
And if you like to avoid doing "chardet" for utf-8, this should work (atleast works for me)
Code:
def getItemDisplayName(itemRef, info, removeExtension=None):
	if itemRef.flags & eServiceReference.isGroup:
		name = itemRef.getName()
	elif itemRef.flags & eServiceReference.isDirectory:
		name = info.getName(itemRef)
		name = path.basename(name.rstrip("/"))
	else:
		name = info.getName(itemRef)
		removeExtension = config.movielist.hide_extensions.value if removeExtension is None else removeExtension
		if removeExtension:
			fileName, fileExtension = path.splitext(name)
			if fileExtension in KNOWN_EXTENSIONS:
				name = fileName
	aname = name.encode('UTF-8', 'surrogateescape')
	if name != aname.decode('UTF-8', 'ignore'):
		from chardet import detect
		encoding = detect(aname)['encoding']
		return aname.decode(encoding)
	return name
 
And if you like to avoid doing "chardet" for utf-8, this should work (atleast works for me)
Code:
def getItemDisplayName(itemRef, info, removeExtension=None):
	if itemRef.flags & eServiceReference.isGroup:
		name = itemRef.getName()
	elif itemRef.flags & eServiceReference.isDirectory:
		name = info.getName(itemRef)
		name = path.basename(name.rstrip("/"))
	else:
		name = info.getName(itemRef)
		removeExtension = config.movielist.hide_extensions.value if removeExtension is None else removeExtension
		if removeExtension:
			fileName, fileExtension = path.splitext(name)
			if fileExtension in KNOWN_EXTENSIONS:
				name = fileName
	aname = name.encode('UTF-8', 'surrogateescape')
	if name != aname.decode('UTF-8', 'ignore'):
		from chardet import detect
		encoding = detect(aname)['encoding']
		return aname.decode(encoding)
	return name

Ah, the joys of Unicode!
 
Currently function getItemDisplayName return value "name" is STR, not bytes.
Yes,
But it should be given a bytes.
So if what is coming in is bytes why does it need to have an encode run on it?
 
And if you like to avoid doing "chardet" for utf-8, this should work (atleast works for me)
Ideally it should work for any filename.
So you can't assume that there is a valid encoding.
You have to get away from the idea that there is some (Windows) language encoding here and handle any arbitrary stream of bytes.
 
@Ocean - still working some of my changes to file_eraser and Trashcan, so have committed your original(working) changes to Git.
We can submit some of your suggestions to reduce code size when this change has been in use for a while and the getItemDisplayName changes added.
 
Last edited:
@ocean, are you able to build or would you like us to supply you with the develop image?
 
@ocean, are you able to build or would you like us to supply you with the develop image?
Thanks, I don't have buildenviroment with me now, but changes looks good and can test tomorrow.

In trashcan change name == ".e2settings.pkl" -> name == b".e2settings.pkl", otherwise comparison can never be true because name is bytes.
 
I think you are talking about a different problem.
I'm talking about post 76 which refers to the code for getItemDisplayName().

name is a filename, so has to be bytes. So why would you run encode on it?

And why, when getting a display name, would you discard information rather than making it visible to the user?
 
I'm talking about post 76 which refers to the code for getItemDisplayName().

name is a filename, so has to be bytes. So why would you run encode on it?

And why, when getting a display name, would you discard information rather than making it visible to the user?

Because external is bytes, but internal is unicode (python 3 automatic conversion unless really forced to bytes … and the python 3 default is always utf-8) ……….. why else would we go through these changes otherwise??
 
Last edited:
Because external is bytes, but internal is unicode (python 3 automatic conversion unless really forced to bytes … and the python 3 default is always utf-8) ……….. why else would we go through these changes otherwise??
The name at this point should be a filename, hence should be bytes.

The filename should be kept as bytes throughout the code (and passed to C+ code as such). Otherwise it cannot cover all possible filenames.
The only time you need it to be anything else is for display, and then the decode should not discard information.
 
The name at this point should be a filename, hence should be bytes.
The filename should be kept as bytes throughout the code (and passed to C+ code as such). Otherwise it cannot cover all possible filenames.
name is str in getItemDisplayName. But what is IMPORTANT here, "filename" is still intact inside str, it's just in different format. This str can always be converted to bytes without any characters lost: name.encode('UTF-8', 'surrogateescape'). Now we have original filename in bytes.
 
Since "getItemDisplayName" is also used in MovieSelection.py for different purposes, maybe it would be best to use two functions instead. This way it's not required to handle possible other codings in MovieSelection.py
Code:
def getItemDisplayName(itemRef, info, removeExtension=None):
	if itemRef.flags & eServiceReference.isGroup:
		name = itemRef.getName()
	elif itemRef.flags & eServiceReference.isDirectory:
		name = info.getName(itemRef)
		name = path.basename(name.rstrip("/"))
	else:
		name = info.getName(itemRef)
		removeExtension = config.movielist.hide_extensions.value if removeExtension is None else removeExtension
		if removeExtension:
			fileName, fileExtension = path.splitext(name)
			if fileExtension in KNOWN_EXTENSIONS:
				name = fileName
	return name


def getItemDisplayNameText(itemRef, info, removeExtension=None):
	name = getItemDisplayName(itemRef, info, removeExtension)
	aname = name.encode('UTF-8', 'surrogateescape')
	if name != aname.decode('UTF-8', 'ignore'):
		from chardet import detect
		encoding = detect(aname)['encoding']
		return aname.decode(encoding)
	return name

And replace all 3 cases of "data.txt = getItemDisplayName(serviceref, info)" with "data.txt = getItemDisplayNameText(serviceref, info)"
 
So with the code changes already made is there still a problem getting these files to play?
 

OpenViX Feeds Status

Back
Top