Superb quality and spec AB-Com PULSe 4K SE. Crazy offer! Only £99! FREE UK DELIVERY! 4K UHD, Enigma 2, Multiboot 4 images & more!...
Superb quality and spec AB-Com PULSe 4K Rev II Twin Satellite tuner only £149! FREE UK DELIVERY! 4K UHD, Enigma 2, SATA HDD facility, Multiboot 4 images & more!...

[OS-mio] Crash when folder contains filenames with Windows encoding

Tested openatv 6.4 and everything works fine there. 0x94 = 148 = ö in codepage 850

Just realised 6.4 still uses python2. Tested Openatv 7.0 and that is also crashing. I guess all pyhton3 images have this issue?!
 
Last edited:
On reflexion 0x94, whilst being perfectly legal in an ext4 filename, is NOT legal utf-8.
That should be the 2-byte sequence 0xc2 0x94.

But ö is U+00f6 in Unicode, which is 0xc3 0xb6 in utf-8.

Anything that looks at filesystem names (or, as in console output, might echo this back) has to be able to cater for the result being non-utf8 when decoded.
So this might be what contributes to triggering the bugs.

I just tried extracting the file on my linux mint system. This how it shows on the console:

Code:
joe@joe-desktop:~/Downloads/mymovies$ ls -al
total 124
drwx------  2 joe joe   4096 Jun 29 09:11  .
drwxr-xr-x 25 joe joe 118784 Jun 29 09:11  ..
-rw-rw-r--  1 joe joe      0 Jun 23 16:44 'p'$'\302\224''ll'$'\302\224''.mp4'

...and in the Nemo File Manager:

Code:
p”ll”.mp4

Where the " are displayed as unicode 0094 symbols.
 
Last edited:
Spent some time debuging code and now know what is causing the crash.

"def createPlaylist()" in file MovieSelection.py

When there is non utf-8 filename in directory, item.getPath() contains surrogates for all non utf-8 characters. It throws exception "std::logic_error" when path is used.

If simply filter out surrogates from item paths like this before appending to items:

item.setPath(item.getPath().encode('UTF-8', 'surrogateescape').decode('UTF-8', 'ignore'))

-> No crash and directory loads normally.

Ofcourse you can't play these files, because path is still incorrect.

How did this work in python27? It should use bytes for file paths, not str.

Anyway, hope this info helps to fix this properly. I did see other similar bug where opening movie directory crashed and it's probably this same issue.
 
How did this work in python27? It should use bytes for file paths, not str.
Agreed. It should. The only time it needs to be "converted" to a string is for display (although even that might not be needed - Py3 can print bytes).
Py2 didn't have the bytes/str distinction, so didn't have this problem.
 
If I may suggest adding temporary workaround for the crash:

Filter out non UTF-8 filenames from movielist. Currently there is no way to play the files anyway.

MovieList.py line 741 has this comment:
# OSX put a lot of stupid files ._* everywhere... we need to skip them

Just add this code after that:

Code:
# Filter out non UTF-8 files. Remove this when movielist supports them.
aname = name.encode('UTF-8', 'surrogateescape').decode('UTF-8', 'ignore')
if aname != name:
    print("[MovieList] skipping non utf-8 filename: %s" % aname)
    continue
 
If I may suggest adding temporary workaround for the crash:

Filter out non UTF-8 filenames from movielist. Currently there is no way to play the files anyway.
The trash handler will still have an issue.
Basically the file-system code needs to use the file-system name when dealing with the file-system. This is a bytes array.
 
The trash handler will still have an issue.
Basically the file-system code needs to use the file-system name when dealing with the file-system. This is a bytes array.
Yes I know, my suggestion only meant to be temporary.

But I think it makes sence, because I don't expect to see fix for filesystem any time soon. Unfortunately.

Same problems are also in other images, so this not VIX only problem. I Wonder if other image developers are aware of this, atleast openatv crashes same way and trashcan code also looks identical.
 
The trash handler will still have an issue.
Basically the file-system code needs to use the file-system name when dealing with the file-system. This is a bytes array.

We are working in python. When we access the file system we use os.walk. Are you saying you think we should remove all python abstraction layers and work directly on the byte code on the hard drive?
 
When we access the file system we use os.walk.

I think using os.walk.. is fine.

In trashcan.py code parameter "trashfolder" is str, that means python3 converts everything to str automatically. Non utf-8 characters are replaced with surrogates -> problem.

Exact same "os.walk" where "trashfolder" parameter is bytes, then root, dirs, files and name are also bytes and no str conversion happens.

That's exactly what we want. But real problem comes from this part: enigma.eBackgroundFileEraser.getInstance().erase(fn)

That function needs accept fn also as bytes.

It's imported from enigma.pyc, but I can't find enigma.py. I guess file is generated somehow from C, is there documentation how this works?

file_eraser.cpp has: void eBackgroundFileEraser::erase(const std::string& filename)

Possible solution is to add overloaded function. I have only minimal experience in C, maybe something like: void eBackgroundFileEraser::erase(const char* filename)

That also likely means whole image needs to be compiled from source before you can even test.
 
I think using os.walk.. is fine.

In trashcan.py code parameter "trashfolder" is str, that means python3 converts everything to str automatically. Non utf-8 characters are replaced with surrogates -> problem.

Exact same "os.walk" where "trashfolder" parameter is bytes, then root, dirs, files and name are also bytes and no str conversion happens.

That's exactly what we want. But real problem comes from this part: enigma.eBackgroundFileEraser.getInstance().erase(fn)

That function needs accept fn also as bytes.

It's imported from enigma.pyc, but I can't find enigma.py. I guess file is generated somehow from C, is there documentation how this works?

file_eraser.cpp has: void eBackgroundFileEraser::erase(const std::string& filename)

Possible solution is to add overloaded function. I have only minimal experience in C, maybe something like: void eBackgroundFileEraser::erase(const char* filename)

That also likely means whole image needs to be compiled from source before you can even test.
Yes it's compiled into Enigma, but all the standard C++ calls in file_eraser such as delete, rename, trunc expect string

building the image is simple, but also somewhat complex.... see bottom of https://github.com/OpenViX/enigma2

appreciate suggestions....... Huevos and I are still looking at having a resolution ... in python
 
Last edited:
appreciate suggestions....... Huevos and I are still looking at having a resolution ... in python
I'm not sure filesystem issues are solvable with ONLY python changes.

In C++ string (std::string) is bytes and it's not required to be UTF-8 string.

C++ functions should be able take bytes, not sure why it's not working. I think if you know answer to this, you would be closed to solution.
 
Last edited:
We are working in python. When we access the file system we use os.walk. Are you saying you think we should remove all python abstraction layers and work directly on the byte code on the hard drive?
No.
I'm saying that when you get a filename (or pathname) you must always keep it as bytes. If you ever convert it to a str then you (may) no longer have anythng that corresponds to what is on disk.
 
I'm not sure filesystem issues are solvable with ONLY python changes.

In C++ string (std::string) is bytes and it's not required to be UTF-8 string.

C++ functions should be able take bytes, not sure why it's not working. I think if you know answer to this, you would be closed to solution.

Can we try experimenting in Trashcan.py?

Code:
				for root, dirs, files in os.walk(trashfolder.encode(), topdown=False):
					for name in files:
						# Don't delete any per-directory config files from .Trash if the option is in use
						if (config.movielist.settings_per_directory.value and name == ".e2settings.pkl"):
							continue
						fn = os.path.join(root, name)
						st = os.stat(fn)
						try:
							import chardet
							encoding = chardet.detect(fn)["encoding"]
							fn = fn.decode(encoding=encoding)
						except Exception as e:
							print( type(e).__name__, e)
						if st.st_ctime < self.ctimeLimit:
							try:
								enigma.eBackgroundFileEraser.getInstance().erase(fn)
								bytesToRemove -= st.st_size
							except Exception as e:
								print("[Trashcan] Failed to erase file after stat.st_ctime selection:", name, "   ", type(e).__name__, e)								
						else:
							candidates.append((st.st_ctime, fn, st.st_size))
							size += st.st_size
					# Remove empty directories if possible
 
Can we try experimenting in Trashcan.py?
It's possible that could work. I can test it later.

Small note, also "name" is bytes. name == ".e2settings.pkl" -> name == b".e2settings.pkl", and also in that debug print
 
Looks like it's still not working. No crash, but it doesn't erase the file.

I modified it to erase always, if filename contains 'test' and added some test files to trash folder.

I guess next step would be to check old python2 image, if it erases the files. (I only know it doesn't crash)

Also if anyone else wants to test, it's easy to create files from terminal. (Just remember to remove st_ctime check from code to make sure it tries to erase)

Code:
fn = b'/media/hdd/movie/.Trash/' + b't\xe4hti.mpg'
f = open(fn, 'w')
f.write('')
f.close()
 
@Ocean - I believe that you are using latin-1?? I tried that with file generated in post #36 and it does same as Ascii.... no delete
So I made the mistake of modifying both file_eraser.cpp and .h to take char*, but image build fails because other C++ modules use this module directly and expect string.
Will look at what it will take to handle this change to include modded file_eraser, but starting to think the whole lot will have to be done in python.
 
@Ocean - I believe that you are using latin-1??
The file-system has no concept of latin-1, unicode, iso8859-16, etc..
It only deals in a byte string.
So what any user is using has to be ignored, except when displaying it to a user, when it might enable it to make sense to a person - but such a translation must not produce an error, which is the problem here.
 
Last edited:
@Ocean,

Is it possible to check with the current code in the repo:
1) it works with utf8 filenames that contain single byte chars > 127.
2) it works with utf8 filenames that contain multi byte chars.
 
So I made the mistake of modifying both file_eraser.cpp and .h to take char*, but image build fails because other C++ modules use this module directly and expect string.
I said it could possible to add overloaded function, not to modify existing one.

I did try compile image from source and there were no errors, but I don't think my changes were included to image. There are no instructions how to compile enigma2 from local folder. Maybe I needed to clean something too. I modified ENIGMA2_URI in site.conf to point to git:///home/.. Could someone add some instructions to github enigma2 page?

My initial cpp and h are attached if you like to test those. Don't know if they compile or not.

Also added debug prints to original function where filename is empty or not found. I think you can commit 2 debug lines if they work, erase shouldn't normally be called with empty or non existing file.

If you get debug prints working you should see in log what filename contains in case it does not erase.
 

Attachments

OpenViX Feeds Status

Back
Top