| ▲ | renjipanicker 3 hours ago | |
Mostly still true, yeah. A few tools have made real progress on error tolerance, continuing past a mistake rather than just stopping, tree-sitter is probably the best example, it's explicitly built to produce a best-effort tree from broken input, which is why it's good for editors. ANTLR also has configurable recovery strategies, single-token insertion/deletion heuristics and the like. But "tolerant" isn't the same as "as good as hand-written." A hand-rolled recursive descent parser can say something like "missing semicolon after return statement" because the code knows exactly what construct it's in. A generated parser's error is usually derived mechanically from the state machine, "expected one of: X, Y, Z, got W", which is correct but generic. This is what Yantra does at the moment. Closing that specific gap would mostly require hand-authored, context-specific messages layered on top. But its a good problem to solve. For yantra specifically, it doesn't have error recovery at all yet. A syntax or lexer error just stops parsing at that point, no resynchronization, no continuing to find more errors in one pass. It's a known, documented gap, not something I'd claim is solved. For the kind of smaller or evolving DSLs this is aimed at, that's probably an acceptable tradeoff, but it does exist as a limitation. | ||
| ▲ | Calavar 2 hours ago | parent | next [-] | |
> A hand-rolled recursive descent parser can say something like "missing semicolon after return statement" because the code knows exactly what construct it's in. > A generated parser's error is usually derived mechanically from the state machine, "expected one of: X, Y, Z, got W", which is correct but generic. You are not describing the difference between hand-written and generated but rather top down vs. bottom up. There may be some confusion here because hand-written parsers are almost always recursive decent (a form of top down), while the most popular parser generators (yacc, bison) are bottom up. However hand-written bottom up parsers do exist (Pratt parsing, recursive ascent), as do top down parser generators (ANTLR, Coco/R, Chumksy). Chumsky in particular gives pretty decent error messages out of the box. Bottom up parsers can parse a larger set of languages than top down parsers, but the tradeoff is you don't know what you are parsing until you successfully reduce the rule (unlike top down parsers). This is why errors in bottom up parsers often lack comments on structure/context like "after return statement" and can only give a list of which alternative symbols would have been valid. | ||
| ▲ | ahmedezat_katte 2 hours ago | parent | prev [-] | |
[flagged] | ||