Re: state in parsers

Kristof Bastiaensen <[email protected]> Tue, 03 Nov 2015 16:57:37 +0100
Newsgroups gmane.emacs.semantic
Message-ID <[email protected]>
Ok, so let's say I emit a 'layout-next' token for implicit ';'.  When 
inside a layout,
this should be part of the grammar rule, and when outside the layout, it 
should
be ignored (like whitespace).  How do I track whether I am inside or 
outside the layout?
Do I need to specify each rule, for example 'expression', twice, one for 
inside, one for
outside?  Or can I set a 'variable' to track this in the parser?

Regards,
Kristof

On 03-11-15 16:28, Eric Ludlam wrote:
>
>
> I am replying on my phone so I'll be a bit brief.
>
> I was providing python as an example of whitespace handling g so you would
> have a starting point.  At a guess, it looks like your keywords like do or
> let also define some sort of indentation depth.  The key for your lexer is
> to create tokens that the parser will use to reason on how it will
> interpret the system.  Your parser would need to accept and use or ignore
> the special tokens depending on context.  I don't think the lexer would
> convert whitespace into braces and ; though if it isn't completely
> deterministic from the lexical point of view.  If you can track what you
> need through the lexer, then you could do the substitution.  The c lexer
> handles preprocessor statements by tracking state while lexing so maybe
> that style will work for you.
>
> Eric
>
> On Nov 3, 2015 9:48 AM, "Kristof Bastiaensen" <[email protected] 
> <mailto:[email protected]>> wrote:
>
>     messed the formatting up again :-( Sorry for the noise...
>
>     Here is a better (contrived but real) example, that demonstrates how
>     layout depends on the parser context:
>     f x = length (
>            do a <- [20]
>               return (30)) + x
>               * 2
>
>     This translates to:
>
>     f x = length (
>            do {a <- [20]
>               ;return (30)}) + x
>               * 2
>
>     Regards,
>     Kristof Bastiaensen
>
>
>     On 03-11-15 15:34, Kristof Bastiaensen wrote:
>     > Actually the first problem can be solved at the lexer, by
>     > remembering the indentation after the 'let', 'where', 'do' or 'of'
>     > keyword.  The second problem is harder, since where to ignore
>     > insert ';' tokens would depend on the parser context.
>     >
>     > The mailer messed up the formatting, so I'll try again,
>     hopefully it is
>     > right now:
>     >
>     > let a = 20
>     >       b = 30
>     >
>     > and
>     > let  a = 20
>     >       b = 30
>     >
>     > are not the same thing.  The first translates to
>     >
>     > let {a = 20
>     >       ;b = 30 }
>     >
>     > and the second to (the illegal)
>     >
>     > let {a = 20}
>     >        b = 30
>     >
>     > And for example the following
>     >
>     > let a = 20
>     >       b = 30
>     >       in a + b
>     >
>     > translates to
>     >
>     > let {a = 20
>     >       ;b = 30
>     >       }in a + b
>     >
>     > not to
>     >
>     > let {a = 30
>     >       ;b = 30
>     >       ;in a + b}
>     >
>     > Which is illegal.  Where to insert ';' depends on the parser
>     context,
>     > it's not possible to do at the lex phase.
>     >
>     > Regards,
>     > Kristof Bastiaensen
>     >
>     >
>     >
>     >
>     >
>     > On 03-11-15 15:21, Kristof Bastiaensen wrote:
>     >> Hi,
>     >>
>     >> the problem in haskell is that simply specifying indent is not
>     enough,
>     >> for example:
>     >>
>     >> let a = 20
>     >>         b = 30
>     >>
>     >> and
>     >> let  a = 20
>     >>         b = 30
>     >>
>     >> are not the same thing.  The first translates to
>     >>
>     >> let {a = 20
>     >>         ;b = 30 }
>     >>
>     >> and the second to (the illegal)
>     >>
>     >> let {a = 20}
>     >>         b = 30
>     >>
>     >> And for example the following
>     >>
>     >> let a = 20
>     >>         b = 30
>     >>         in a + b
>     >>
>     >> translates to
>     >>
>     >> let {a = 20
>     >>         ;b = 30
>     >>         }in a + b
>     >>
>     >> not to
>     >>
>     >> let {a = 30
>     >>         ;b = 30
>     >>         ;in a + b}
>     >>
>     >> Which is illegal.  Where to insert ';' depends on the parser
>     really,
>     >> it's hard to do at the lex phase.
>     >>
>     >> Regards,
>     >> Kristof Bastiaensen
>     >>
>     >>
>     >>
>     >> On 03-11-15 13:22, Eric Ludlam wrote:
>     >>> On 11/02/2015 07:13 AM, Kristof Bastiaensen wrote:
>     >>>> Hi,
>     >>>>
>     >>>> I'd like to add support for haskell to semantic.  Haskell
>     syntax is
>     >>>> sensitive
>     >>>> to indentation, and I'd like to know if it is possible to use
>     bovine or
>     >>>> wisent for
>     >>>> it.
>     >>> Yes.
>     >>>
>     >>>> The layout rule is explained here:
>     >>>>
>     https://www.haskell.org/onlinereport/haskell2010/haskellch2.html#x7-210002.7
>     >>>>
>     >>>>
>     >>>> What it does, is after certain keywords, whenever a '{'
>     character is
>     >>>> ommited, it
>     >>>> enters a layout context.  In the layout context it will
>     remember the
>     >>>> indentation
>     >>>> of the next token.  When a line has the same indentation as
>     this token,
>     >>>> a ';' token is
>     >>>> inserted before the expression.  Whenever the indentation is
>     less than
>     >>>> this, a '}' token
>     >>>> is inserted, and the layout context is ended.  The layout
>     context is
>     >>>> also ended when a parser error would occur.
>     >>> The python grammar in semantic/wisent/python.wy and python.el
>     is an ok
>     >>> example.  The premise is that when you develop the lexer, you
>     teach it
>     >>> to convert your indentation into lexical tokens that you then
>     use in
>     >>> your grammar.
>     >>>
>     >>> In the python.el support file, you will notice it uses
>     >>> (current-indentation) and uses that to derive INDENT and DEDENT
>     >>> tokens.  The grammar uses those to create INDENT and DEDENT
>     blocks so
>     >>> it can recurse into code bodies.  You could just as easily
>     skip them
>     >>> too if you prefer.
>     >>>
>     >>>> So I would set a state variable to all the layout context
>     indentations,
>     >>>> and then change the behaviour of the lexer or parser based on
>     this
>     >>>> state.  Is that possible?  Would it interfere with
>     incremental parsing?
>     >>> Definitely start with the lexer, and there is good
>     documentation on
>     >>> creating a lexer in the semantic language developers guide.
>     Use the
>     >>> lexer testing function to make sure it outputs what you
>     expect, and
>     >>> then you can continue on with the grammar.
>     >>>
>     >>> Eric
>     >>
>     ------------------------------------------------------------------------------
>     >> _______________________________________________
>     >> cedet-semantic mailing list
>     >> [email protected]
>     <mailto:[email protected]>
>     >> https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>     >
>     >
>     ------------------------------------------------------------------------------
>     > _______________________________________________
>     > cedet-semantic mailing list
>     > [email protected]
>     <mailto:[email protected]>
>     > https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>
>
>     ------------------------------------------------------------------------------
>     _______________________________________________
>     cedet-semantic mailing list
>     [email protected]
>     <mailto:[email protected]>
>     https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>

------------------------------------------------------------------------------

_______________________________________________
cedet-semantic mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/cedet-semantic