Summary: The current description of the WGSL tokenization and parser is not accurate. Something more subtle is needed.
The WGSL spec describes tokenization as independent of parsing.
But Treesitter, which we use to check for validity and potential ambiguity, uses a few tricks that go beyond textbook parsing.
See https://tree-sitter.github.io/tree-sitter/creating-parsers#conflicting-tokens
The first thing mentioned there, context-aware lexing, is essential to making the WGSL grammar "work".
The background research is referenced from the Treesitter docs:
Context-Aware Scanning for Parsing Extensible Languages Eric R. Van Wyk, August C. Schwerdfeger, GCPE'07, 2007.
The key idea is there's cooperation between the parser and the tokenizer. When asking for a token, the parser tells the token-scanner which tokens are acceptable to return, assuming the current parser state. If your grammar is reasonable, then it will all be deterministic and predictable.
For example, consider let v: vec2<i32>= vec2(1,2);
The two code points > = are adjacent. The greedy tokenization, as described by the WGSL spec would yield the greater_than_equal token, and cause the parse to fail.
Instead, Treesitter knows that at that point in the parse, immediately after the i32 token, that > is an acceptable next token and >= is not. So when it requests a next token, the scanner is given that discriminating info, and the scanner will return the > token. Then the parse can succeed.
Hand-written parsers, like the ones in Tint, handle this with a manually-specified retry. First it tries to consume the greater_than_equal token, and if the parse would fail, it backs of and splits the token into two.
Note: I found this when exploring how to add the shift tokens back again (due to popular demand).
Summary: The current description of the WGSL tokenization and parser is not accurate. Something more subtle is needed.
The WGSL spec describes tokenization as independent of parsing.
But Treesitter, which we use to check for validity and potential ambiguity, uses a few tricks that go beyond textbook parsing.
See https://tree-sitter.github.io/tree-sitter/creating-parsers#conflicting-tokens
The first thing mentioned there, context-aware lexing, is essential to making the WGSL grammar "work".
The background research is referenced from the Treesitter docs:
Context-Aware Scanning for Parsing Extensible Languages Eric R. Van Wyk, August C. Schwerdfeger, GCPE'07, 2007.
The key idea is there's cooperation between the parser and the tokenizer. When asking for a token, the parser tells the token-scanner which tokens are acceptable to return, assuming the current parser state. If your grammar is reasonable, then it will all be deterministic and predictable.
For example, consider
let v: vec2<i32>= vec2(1,2);The two code points
>=are adjacent. The greedy tokenization, as described by the WGSL spec would yield thegreater_than_equaltoken, and cause the parse to fail.Instead, Treesitter knows that at that point in the parse, immediately after the
i32token, that>is an acceptable next token and>=is not. So when it requests a next token, the scanner is given that discriminating info, and the scanner will return the>token. Then the parse can succeed.Hand-written parsers, like the ones in Tint, handle this with a manually-specified retry. First it tries to consume the greater_than_equal token, and if the parse would fail, it backs of and splits the token into two.
Note: I found this when exploring how to add the shift tokens back again (due to popular demand).