Skip to content

Consider emitting buffered DEDENT tokens on the last line #104976

Description

@pablogsal

In Python 3.12, porting the tokenizer to use the C tokenizer underneath to support PEP 701 has now a documented change in docs.python.org/3.12/whatsnew/3.12.html#changes-in-the-python-api:

Some final DEDENT tokens are now emitted within the bounds of the input. This means that for a file containing 3 lines, the old version of the tokenizer returned a DEDENT token in line 4 whilst the new version returns the token in line 3.

Apparently, this affects negatively some formatting tools (see PyCQA/pycodestyle#1142). Let's consider what options do we have and see if we can fix this without adding a lot of maintenance burden to the C tokenizer or slowing down everything.

Linked PRs

Activity

  1. pablogsal commented on May 26, 2023

    @pablogsal
    MemberAuthor

    We can do something like this:

    diff --git a/Lib/tokenize.py b/Lib/tokenize.py
    index 911f0f12f9..63dc44b0dc 100644
    --- a/Lib/tokenize.py
    +++ b/Lib/tokenize.py
    @@ -452,7 +452,9 @@ def _tokenize(rl_gen, encoding):
             yield token
         if token is not None:
             last_line, _ = token.start
    -        yield TokenInfo(ENDMARKER, '', (last_line + 1, 0), (last_line + 1, 0), '')
    +        if token.type != DEDENT:
    +            last_line != 1
    +        yield TokenInfo(ENDMARKER, '', (last_line, 0), (last_line, 0), '')
    
    
     def generate_tokens(readline):
    diff --git a/Python/Python-tokenize.c b/Python/Python-tokenize.c
    index 88087c1256..2a09dfd94a 100644
    --- a/Python/Python-tokenize.c
    +++ b/Python/Python-tokenize.c
    @@ -214,6 +214,10 @@ tokenizeriter_next(tokenizeriterobject *it)
         }
    
         if (it->tok->tok_extra_tokens) {
    +        if (type == DEDENT && it->tok->done == E_EOF) {
    +            lineno = end_lineno = lineno + 1;
    +            col_offset = end_col_offset = 0;
    +        }
             // Necessary adjustments to match the original Python tokenize
             // implementation
             if (type > DEDENT && type < OP) {

    but I think this forces us to somehow handle the ENDMARKER internally. Maybe that's a possible solution but I fear this still has some side effects.

  2. pablogsal commented on May 26, 2023

    @pablogsal
    MemberAuthor

    @lysnikolaou thoughts?

  3. added a commit that references this issue on May 26, 2023
  4. pablogsal commented on May 26, 2023

    @pablogsal
    MemberAuthor
  5. self-assigned this
    on May 26, 2023
  6. pablogsal commented on May 26, 2023

    @pablogsal
    MemberAuthor

    Opened #104980 to test this idea

  7. lysnikolaou commented on May 26, 2023

    @lysnikolaou
    Member

    Yeah, I think that, if we want to support doing the same thing as 3.11, the only way is to special-case it Python-tokenize.c and not in the C tokenizer itself.

  8. added 3 commits that reference this issue on May 26, 2023
  9. pablogsal commented on May 26, 2023

    @pablogsal
    MemberAuthor

    Yeah, I think that, if we want to support doing the same thing as 3.11, the only way is to special-case it Python-tokenize.c and not in the C tokenizer itself.

    Ok, then check if you like #104980

  10. added 4 commits that reference this issue on May 26, 2023
  11. added a commit that references this issue on May 26, 2023
  12. added 2 commits that reference this issue on May 26, 2023
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

No labels
No labels

Projects

No projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions