| ▲ | traes 3 hours ago | |||||||||||||
I recall reading in the past that the primary reason parser generators aren't used for production compilers is the difficulty of making them produce useful error messages on malformed code. (Of course, the need for fine tuned optimizations also plays a role). Is this still true, or have parser generators caught up in this regard? | ||||||||||||||
| ▲ | renjipanicker 3 hours ago | parent [-] | |||||||||||||
Mostly still true, yeah. A few tools have made real progress on error tolerance, continuing past a mistake rather than just stopping, tree-sitter is probably the best example, it's explicitly built to produce a best-effort tree from broken input, which is why it's good for editors. ANTLR also has configurable recovery strategies, single-token insertion/deletion heuristics and the like. But "tolerant" isn't the same as "as good as hand-written." A hand-rolled recursive descent parser can say something like "missing semicolon after return statement" because the code knows exactly what construct it's in. A generated parser's error is usually derived mechanically from the state machine, "expected one of: X, Y, Z, got W", which is correct but generic. This is what Yantra does at the moment. Closing that specific gap would mostly require hand-authored, context-specific messages layered on top. But its a good problem to solve. For yantra specifically, it doesn't have error recovery at all yet. A syntax or lexer error just stops parsing at that point, no resynchronization, no continuing to find more errors in one pass. It's a known, documented gap, not something I'd claim is solved. For the kind of smaller or evolving DSLs this is aimed at, that's probably an acceptable tradeoff, but it does exist as a limitation. | ||||||||||||||
| ||||||||||||||