Sitelet https://github.com/python/cpython/issues/85308
Skip to content

argparse uses default encoding when read arguments from file #85308

Description

@serhiy-storchaka
BPO 41136
Nosy @rhettinger, @serhiy-storchaka, @ZackerySpytz
PRs
  • gh-85308: Add a parameter to argparse.ArgumentParser for file encoding #21444
  • Note: these values reflect the state of the issue at the time it was migrated and might not reflect the current state.

    Show more details

    GitHub fields:

    assignee = None
    closed_at = None
    created_at = <Date 2020-06-27.07:13:47.818>
    labels = ['type-feature', 'library', '3.10']
    title = 'argparse uses default encoding when read arguments from file'
    updated_at = <Date 2020-07-11.12:13:17.671>
    user = 'https://github.com/serhiy-storchaka'

    bugs.python.org fields:

    activity = <Date 2020-07-11.12:13:17.671>
    actor = 'ZackerySpytz'
    assignee = 'none'
    closed = False
    closed_date = None
    closer = None
    components = ['Library (Lib)']
    creation = <Date 2020-06-27.07:13:47.818>
    creator = 'serhiy.storchaka'
    dependencies = []
    files = []
    hgrepos = []
    issue_num = 41136
    keywords = ['patch']
    message_count = 1.0
    messages = ['372452']
    nosy_count = 4.0
    nosy_names = ['rhettinger', 'paul.j3', 'serhiy.storchaka', 'ZackerySpytz']
    pr_nums = ['21444']
    priority = 'normal'
    resolution = None
    stage = 'patch review'
    status = 'open'
    superseder = None
    type = 'enhancement'
    url = 'https://bugs.python.org/issue41136'
    versions = ['Python 3.10']

    Activity

    1. serhiy-storchaka commented on Jun 27, 2020

      @serhiy-storchaka
      MemberAuthor

      The fromfile_prefix_chars option allows you to read arguments from file. But open() without explicit encoding is used for this. Therefore the result is depending on the current locale or the PYTHONIOENCODING environment variable. On Linux this is rare a problem, because UTF-8 is a common locale encoding, but on Windows the locale encoding is usually 8-bit and may even be not able to encode all file names.

      I think we need a new option to specify the encoding for files from which arguments are read.

    2. added
      stdlibStandard Library Python modules in the Lib/ directory
      type-featureA feature request or enhancement
      on Jun 27, 2020
    3. transferred this issue fromon Apr 10, 2022
    4. methane commented on May 25, 2022

      @methane
      Member

      I think sys.getfilesystemencoding() and sys.getfilesystemencodeerrors() should be default.

      I am not sure encoding/errors option is worth enough.
      On Unix, fsencoding is locale encoding on Unix. UTF-8 mode affects both.
      On Windows, PYTHONLEGACYWINDOWSFSENCODING can be used for backward compat option.

      @vstinner How do you think?

    5. vstinner commented on May 26, 2022

      @vstinner
      Member

      argparse has a big API, which specific API are you talking about? type=argparse.FileType accepts an encoding:

      parser.add_argument('dest_file', type=argparse.FileType('w', encoding='latin-1'))

      The documentation gives this example:

      parser.add_argument('source_file', type=open)

      You can use functools.partial(open, encoding='utf-8') to pass an encoding if you really don't want to use FileType.

    6. vstinner commented on May 26, 2022

      @vstinner
      Member

      I expect that FileType() used with no encoding behaves the same than open() used with no encoding.

    7. methane commented on May 27, 2022

      @methane
      Member

      @vstinner We are talking about here.

      cpython/Lib/argparse.py

      Lines 2164 to 2169 in 20d30ba

      with open(arg_string[1:]) as args_file:
      arg_strings = []
      for arg_line in args_file.read().splitlines():
      for arg in self.convert_arg_line_to_args(arg_line):
      arg_strings.append(arg)
      arg_strings = self._read_args_from_files(arg_strings)

      argparse reads options from file with default text encoding for now.

      On Unix, commandline arguments are decoded with fsencoding. So fsencoding should be used to parse option file too. This is almost backward compatible because locale encoding is same to fsencoding in most cases.

      On Windows, fsencoding is UTF-8 by default and can be ANSI code page when PYTHONLEGACYWINDOWSFSENCODING is set. This is slightly backward incompatible, but I think it is perfect choice because:

      • Commandline arguments and filesystem encoding are UTF-16 (WTF-16). UTF-8/surrogatepass is the best text encoding to represent it.
      • Notepad.exe and most major text editors uses UTF-8 by default.
      • User can set PYTHONLEGACYWINDOWSFSENCODING if they have option file with ANSI code page and they can not convert it to UTF-8.

      So I don't think we need to add another option like fromfile_encoding in #21444.

    8. added a commit that references this issue on May 27, 2022
    9. 2 remaining items

    10. methane commented on May 27, 2022

      @methane
      Member

      I understand that argparse uses UTF-8 uses most operating systems, but the ANSI code page on Windows. If you want to change something, why not using UTF-8 on all platforms?

      Because arguments file should be consistent with commandline arguments on Unix.
      When we make UTF-8 mode default, they are UTF-8.

      I would prefer using the strict error handler. If a file is not encoded to UTF-8, the developer should fix their code. From example, if a file is encoded to Latin1 or a Windows ANSI code page, using UTF-8/surrogateescape would produce a lot of surrogate characters. On Windows, I don't understand well why surrogatepass should be used.

      Because files and command line arguments on Windows may be not valid UTF-16. (WTF-16).
      PEP 529 describes why surrogatepass is chosen.

      I agree that using the filesystem encoding is reasonable on Linux. Are there any options/envvars which affect the encoding of command args? Are there any options/envvars which affect the filesystem encoding, but not the encoding of command args?

      On Unix, fsencoding is immutable and it must be used to decode commandline. UTF-8 Mode affects it.

      On Windows, PYTHONLEGACYWINDOWSFSENCODING and sys._enablelegacywindowsfsencoding() change the fsencoding.
      Commandline arguments are passed by wchar_t* so fsencoding is not used directly.

      But UTF-8/surrogatepass (e.g. WTF-8) is the best way to hold wchar_t* path and commandlines.

      What encoding is used on Windows for batch files? What encoding is used in dir >files. @eryksun

      On cmd.exe, console codepage is used.
      On modern Powershell, $OutputEncoding is used and it is UTF-8 by default.
      On cygwin and msys, ls > files is UTF-8.

    11. vstinner commented on May 27, 2022

      @vstinner
      Member

      I dislike sys._enablelegacywindowsfsencoding() function. For me, it's a bad idea to change the encoding at runtime. Strings already decoded from or encode to UTF-8 will cause mojibake once Python switchs to the ANSI code page.

      I would also prefer to deprecate PYTHONLEGACYWINDOWSFSENCODING. I don't know any issue. It seems like the migration to UTF-8 was just smooth on Windows, since Python hides the encoding and just expose Unicode for most APIs.

    12. vstinner commented on May 27, 2022

      @vstinner
      Member

      Because arguments file should be consistent with commandline arguments on Unix.

      Sadly, it knows it too well that attempting to have two encodings for two unrelated things usually end with mojibake since unrelated things are used together in some cases. For example, filenames can be passed on the command line and in an environment variable. All attempts to "use two encodings" failed badly: https://vstinner.github.io/painful-history-python-filesystem-encoding.html

      I'm fine with using the filesystem encoding for this argparse function.

    13. methane commented on May 27, 2022

      @methane
      Member

      Mercurial has sad history and still relying on sys._enablelegacywindowsfsencoding().

      They had plan to go to UTF-8. https://www.mercurial-scm.org/wiki/WindowsUTF8Plan
      But I don't know about they are working on the UTF-8 plan.

      AFAIk, Mercurial team are planning to rewrite Mercurial in Rust with embedded Python.
      And Rust also uses WTF-8 (UTF-8/surrogatepass) for paths and commandlines on Windows.

    14. eryksun commented on May 27, 2022

      @eryksun
      Contributor

      @serhiy-storchaka, CMD writes to disk and pipe files using the console output code page, except it uses UTF-16 if it's started with the /u command-line option. The default console output code page is set to the system OEM code page, but it can be configured differently in the registry, or set dynamically via "chcp.com" or WinAPI SetConsoleOutputCP(). The system OEM code page is usually the legacy OEM code page of the system locale (e.g. 437, 850), but in Windows 10+ the system OEM code page can be set to UTF-8 (65001).

    15. Repository owner moved this from Features to Doc issues in Argparse issueson Jun 23, 2022
    16. added a commit that references this issue on Jun 23, 2022
    17. added 3 commits that reference this issue on Jun 23, 2022
    18. moved this from Doc issues to Done in Argparse issueson Jul 30, 2022
    Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

    Metadata

    Metadata

    Assignees

    No one assigned

      Labels

      stdlibStandard Library Python modules in the Lib/ directorytype-featureA feature request or enhancement

      Projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions