Repository navigation
argparse uses default encoding when read arguments from file #85308
Description
Activity
The fromfile_prefix_chars option allows you to read arguments from file. But open() without explicit encoding is used for this. Therefore the result is depending on the current locale or the PYTHONIOENCODING environment variable. On Linux this is rare a problem, because UTF-8 is a common locale encoding, but on Windows the locale encoding is usually 8-bit and may even be not able to encode all file names.
I think we need a new option to specify the encoding for files from which arguments are read.
- added3.10 (EOL)end of lifeend of lifestdlibStandard Library Python modules in the Lib/ directoryStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancementA feature request or enhancement
on Jun 27, 2020 I think
sys.getfilesystemencoding()andsys.getfilesystemencodeerrors()should be default.I am not sure encoding/errors option is worth enough.
On Unix, fsencoding is locale encoding on Unix. UTF-8 mode affects both.
On Windows, PYTHONLEGACYWINDOWSFSENCODING can be used for backward compat option.@vstinner How do you think?
argparse has a big API, which specific API are you talking about?
type=argparse.FileTypeaccepts an encoding:parser.add_argument('dest_file', type=argparse.FileType('w', encoding='latin-1'))The documentation gives this example:
parser.add_argument('source_file', type=open)You can use
functools.partial(open, encoding='utf-8')to pass an encoding if you really don't want to use FileType.I expect that FileType() used with no encoding behaves the same than open() used with no encoding.
@vstinner We are talking about here.
Lines 2164 to 2169 in 20d30ba
with open(arg_string[1:]) as args_file: arg_strings = [] for arg_line in args_file.read().splitlines(): for arg in self.convert_arg_line_to_args(arg_line): arg_strings.append(arg) arg_strings = self._read_args_from_files(arg_strings) argparse reads options from file with default text encoding for now.
On Unix, commandline arguments are decoded with fsencoding. So fsencoding should be used to parse option file too. This is almost backward compatible because locale encoding is same to fsencoding in most cases.
On Windows, fsencoding is UTF-8 by default and can be ANSI code page when PYTHONLEGACYWINDOWSFSENCODING is set. This is slightly backward incompatible, but I think it is perfect choice because:
- Commandline arguments and filesystem encoding are UTF-16 (WTF-16). UTF-8/surrogatepass is the best text encoding to represent it.
- Notepad.exe and most major text editors uses UTF-8 by default.
- User can set PYTHONLEGACYWINDOWSFSENCODING if they have option file with ANSI code page and they can not convert it to UTF-8.
So I don't think we need to add another option like
fromfile_encodingin #21444.- added a commit that references this issue
on May 27, 2022 2 remaining items
I understand that argparse uses UTF-8 uses most operating systems, but the ANSI code page on Windows. If you want to change something, why not using UTF-8 on all platforms?
Because arguments file should be consistent with commandline arguments on Unix.
When we make UTF-8 mode default, they are UTF-8.I would prefer using the strict error handler. If a file is not encoded to UTF-8, the developer should fix their code. From example, if a file is encoded to Latin1 or a Windows ANSI code page, using UTF-8/surrogateescape would produce a lot of surrogate characters. On Windows, I don't understand well why surrogatepass should be used.
Because files and command line arguments on Windows may be not valid UTF-16. (WTF-16).
PEP 529 describes why surrogatepass is chosen.I agree that using the filesystem encoding is reasonable on Linux. Are there any options/envvars which affect the encoding of command args? Are there any options/envvars which affect the filesystem encoding, but not the encoding of command args?
On Unix, fsencoding is immutable and it must be used to decode commandline. UTF-8 Mode affects it.
On Windows, PYTHONLEGACYWINDOWSFSENCODING and
sys._enablelegacywindowsfsencoding()change the fsencoding.
Commandline arguments are passed bywchar_t*so fsencoding is not used directly.But UTF-8/surrogatepass (e.g. WTF-8) is the best way to hold
wchar_t*path and commandlines.What encoding is used on Windows for batch files? What encoding is used in
dir >files. @eryksunOn cmd.exe, console codepage is used.
On modern Powershell,$OutputEncodingis used and it is UTF-8 by default.
On cygwin and msys,ls > filesis UTF-8.I dislike
sys._enablelegacywindowsfsencoding()function. For me, it's a bad idea to change the encoding at runtime. Strings already decoded from or encode to UTF-8 will cause mojibake once Python switchs to the ANSI code page.I would also prefer to deprecate PYTHONLEGACYWINDOWSFSENCODING. I don't know any issue. It seems like the migration to UTF-8 was just smooth on Windows, since Python hides the encoding and just expose Unicode for most APIs.
Because arguments file should be consistent with commandline arguments on Unix.
Sadly, it knows it too well that attempting to have two encodings for two unrelated things usually end with mojibake since unrelated things are used together in some cases. For example, filenames can be passed on the command line and in an environment variable. All attempts to "use two encodings" failed badly: https://vstinner.github.io/painful-history-python-filesystem-encoding.html
I'm fine with using the filesystem encoding for this argparse function.
Mercurial has sad history and still relying on
sys._enablelegacywindowsfsencoding().They had plan to go to UTF-8. https://www.mercurial-scm.org/wiki/WindowsUTF8Plan
But I don't know about they are working on the UTF-8 plan.AFAIk, Mercurial team are planning to rewrite Mercurial in Rust with embedded Python.
And Rust also uses WTF-8 (UTF-8/surrogatepass) for paths and commandlines on Windows.@serhiy-storchaka, CMD writes to disk and pipe files using the console output code page, except it uses UTF-16 if it's started with the
/ucommand-line option. The default console output code page is set to the system OEM code page, but it can be configured differently in the registry, or set dynamically via "chcp.com" or WinAPISetConsoleOutputCP(). The system OEM code page is usually the legacy OEM code page of the system locale (e.g. 437, 850), but in Windows 10+ the system OEM code page can be set to UTF-8 (65001).
Metadata
Metadata
Assignees
Labels
Projects
- StatusShow more project fieldsDone
Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.
Show more details
GitHub fields:
bugs.python.org fields: