Repository navigation
Bytecode positions seem way too broad #93691
Description
Activity
- addedtype-bugAn unexpected behavior, bug, or errorAn unexpected behavior, bug, or errorinterpreter-core(Objects, Python, Grammar, and Parser dirs)(Objects, Python, Grammar, and Parser dirs)3.11only security fixesonly security fixes3.12only security fixesonly security fixes
on Jun 10, 2022 I can perhaps see why the argument could be made that we should have the location info for certain constructs span their entire block or jump range, but to me this just feels like the shape of the AST and the design of the compiler are leaking into the bytecode more than is really helpful in practice.
As a quick-and-dirty experiment: for the given example function, setting
node.end_lineno = node.linenoon every node before compiling the AST resulted in a 20% reduction in the size ofco_linetable.I can perhaps see why the argument could be made that we should have the location info for certain constructs span their entire block or jump range, but to me this just feels like the shape of the AST and the design of the compiler are leaking into the bytecode more than is really helpful in practice.
I don't think that many of these things were conscious decisions. Originally we added and enabled the infrastructure so the debug information could be propagated and we spent some time doing small optimisations, but we are missing a full pass over the compiler to fix things like this. Additionally, there are many instructions that don't really benefit from having position information because they are either artificial or don't map well to source code.
I think this is a very good find. With what seems like a small tedious amount of work we could reduce substantially the size for some functions, specially in block setup stuff.
Thanks for opening the issue and the insights @brandtbucher, this is very interesting indeed.
I don't think the shape of the AST is "leaking" into the bytecode. The AST defines the locations.
The problem is, IMO, in the design of the compiler. Tracking the "current" location in the compiler only makes sense if the bytecode is produced in a linear fashion, which it clearly isn't.We should make the location used explicit when generating code, not use the implicit location stored in the compiler.
ADDOP(opcode, oparg,location)6 remaining items
- added a commit that references this issue
on Jun 30, 2024
(Note that
discurrently has a bug in displaying accurate location info in the presence ofCACHEs. The correct information can be observed by working withco_positionsdirectly or using the code from that PR.)While developing
specialist, I realized that there are lots of common code patterns that produce bytecode with unexpectedly large source ranges. In addition to being unhelpful for both friendly tracebacks (the original motivation) and things like bytecode introspection, I suspect these huge ranges may also be bloating the size of our internal position tables as well.Consider the following function:
Things that should probably span one line at most:
GET_ITER/FOR_ITERpair span all of lines 4 through 10.GET_ITER/FOR_ITERpair spans all of lines 5 through 10.POP_JUMP_FORWARD_IF_FALSEspans all of lines 6 through 9.POP_JUMP_FORWARD_IF_FALSEspans all of lines 8 through 9.withcleanup each span all of lines 3 through 10.Things that should probably be artificial:
JUMP_FORWARDspans all of line 7.JUMP_BACKWARDspans all of line 10.JUMP_BACKWARDspans all of lines 5 through 10.Things I don't get:
NOPspans all of lines 4 through 10.As a result, over half of the generated bytecode for this function claims to span line 9, for instance. Also not shown here: the instructions for building functions and classes have similarly huge spans.
I think this can be tightened up in the compiler by:
SET_LOCon child nodes.UNSET_LOCbefore unconditional jumps.Linked PRs