Sitelet https://github.com/python/cpython/issues/105017
Skip to content

Tokenizer produces different output on Windows on py312 for ends of files #105017

Description

@AlexWaygood

Bug report

If you copy and paste the following code into a repro.py file and run python -m tokenize on it on a Windows machine, the output is different on 3.12/3.13 compared to what it was on 3.11 (the file ends with a single newline):

foo = 'bar'
spam = 'eggs'

On Python 3.11, on Windows, the output is this:

> python -m tokenize cpython/repro.py
0,0-0,0:            ENCODING       'utf-8'
1,0-1,3:            NAME           'foo'
1,4-1,5:            OP             '='
1,6-1,11:           STRING         "'bar'"
1,11-1,13:          NEWLINE        '\r\n'
2,0-2,4:            NAME           'spam'
2,5-2,6:            OP             '='
2,7-2,13:           STRING         "'eggs'"
2,13-2,15:          NEWLINE        '\r\n'
3,0-3,0:            ENDMARKER      ''

On Python 3.13 (@ 6e62eb2) on Windows, however, the output is this:

> python -m tokenize repro.py
0,0-0,0:            ENCODING       'utf-8'
1,0-1,3:            NAME           'foo'
1,4-1,5:            OP             '='
1,6-1,11:           STRING         "'bar'"
1,11-1,12:          NEWLINE        '\n'
2,0-2,4:            NAME           'spam'
2,5-2,6:            OP             '='
2,7-2,13:           STRING         "'eggs'"
2,13-2,14:          NEWLINE        '\n'
3,0-3,1:            NL             '\n'
4,0-4,0:            ENDMARKER      ''

There appear to be two changes here:

  • All the NEWLINE tokens now have \n values, whereas on Python 3.11 they all had \r\n values
  • There is an additional NL token at the end, immediately before the ENDMARKER token.

As discussed in PyCQA/pycodestyle#1142, this appears to be Windows-specific, and may be the cause of a large number of spurious W391 errors from the pycodestyle linting tool. (W391 dictates that there should be one, and only one, newline at the end of a file.) The pycodestyle tool is included in flake8, meaning that test failures in pycodestyle can cause test failures for other flake8 plugins. (All tests for the flake8-pyi plugin, for example, currently fail on Python 3.13 on Windows.)

Your environment

Python 3.13.0a0 (heads/main:6e62eb2e70, May 27 2023, 14:00:13) [MSC v.1932 64 bit (AMD64)] on win32

Linked PRs

Activity

  1. added
    type-bugAn unexpected behavior, bug, or error
    3.12only security fixes
    3.13only security fixes
    on May 27, 2023
  2. changed the title [-]Tokenizer produces different output on Windows for ends of files[/-] [+]Tokenizer produces different output on Windows on py312 for ends of files[/+] on May 27, 2023
  3. pablogsal commented on May 27, 2023

    @pablogsal
    Member
  4. AlexWaygood commented on May 27, 2023

    @AlexWaygood
    MemberAuthor

    #105022 fixes the issue of the additional NL token at the end of the file on Windows, but there are still differences on Windows between what the tokenizer produced on 3.11 and what it produces on #105022.

    On 3.11 on Windows:

    >python -m tokenize cpython/repro.py
    0,0-0,0:            ENCODING       'utf-8'
    1,0-1,3:            NAME           'foo'
    1,4-1,5:            OP             '='
    1,6-1,11:           STRING         "'bar'"
    1,11-1,13:          NEWLINE        '\r\n'
    2,0-2,4:            NAME           'spam'
    2,5-2,6:            OP             '='
    2,7-2,13:           STRING         "'eggs'"
    2,13-2,15:          NEWLINE        '\r\n'
    3,0-3,0:            ENDMARKER      ''
    

    Using #105022:

    > python -m tokenize repro.py
    0,0-0,0:            ENCODING       'utf-8'
    1,0-1,3:            NAME           'foo'
    1,4-1,5:            OP             '='
    1,6-1,11:           STRING         "'bar'"
    1,11-1,12:          NEWLINE        '\n'
    2,0-2,4:            NAME           'spam'
    2,5-2,6:            OP             '='
    2,7-2,13:           STRING         "'eggs'"
    2,13-2,14:          NEWLINE        '\n'
    3,0-3,0:            ENDMARKER      ''
    

    Note that the same tokens are emitted now as were emitted on 3.11, but the end-column coordinate is off by one for each NEWLINE token.

    (I'm still seeing loads of flake8-pyi test failures even using #105022, and I'm presuming this is the cause?)

  5. added a commit that references this issue on May 27, 2023
  6. added a commit that references this issue on May 27, 2023
  7. added a commit that references this issue on May 27, 2023
  8. pablogsal commented on May 28, 2023

    @pablogsal
    Member

    @AlexWaygood can you check again with the new PR (#105030)?

  9. added a commit that references this issue on May 28, 2023
  10. added a commit that references this issue on May 28, 2023
  11. added a commit that references this issue on May 28, 2023
  12. AlexWaygood commented on May 28, 2023

    @AlexWaygood
    MemberAuthor

    @AlexWaygood can you check again with the new PR (#105030)?

    Will check either this evening or tomorrow

  13. AlexWaygood commented on May 28, 2023

    @AlexWaygood
    MemberAuthor

    Well, I have good news and bad news.

    The good news is that, on Windows, the tokenize module appears to now be producing exactly the same output on main as it did on 3.11.

    The bad news is that the spurious pycodestyle W391 errors discussed in PyCQA/pycodestyle#1142 still haven't gone away, even using CPython main. These definitely bisect to 6715f91, so something else must be going on to cause these failures :(

    Regardless, the original bug reported here has definitely been fixed, so I'll close this for now, and I'll open a new issue once I've done some more investigation into what might be causing the spurious W391 errors and (hopefully) found a minimal repro that doesn't involve third-party code. Thanks @pablogsal and @mgmacias95 for all your hard work on this!

  14. AlexWaygood commented on May 28, 2023

    @AlexWaygood
    MemberAuthor

    I'll open a new issue once I've done some more investigation into what might be causing the spurious W391 errors and (hopefully) found a minimal repro that doesn't involve third-party code. Thanks @pablogsal and @mgmacias95 for all your hard work on this!

    I've tried to investigate, but I can't immediately figure out what might be causing these W391 errors. I suspect it'll need somebody more familiar with pycodestyle's architecture and/or the tokenize module to figure out whether this is, in fact, a CPython bug at all -- and if it is, to produce a minimal repro that doesn't depend on pycodestyle.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Labels

3.12only security fixes3.13only security fixesOS-windowstype-bugAn unexpected behavior, bug, or error

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions