Re: {Spam?} Re: state in parsers

Eric Ludlam <[email protected]> Fri, 20 Nov 2015 21:29:59 -0500
Newsgroups gmane.emacs.semantic
Message-ID <[email protected]>
Sorry for the delay in replying.  Kids take up a lot of time.

Based on your more recent example, I think your lexer can do the 
tracking.  In this case, the lexer probably needs to track several things.

1) when it sees "do" it needs to create "do" and "{"
2) when it sees ")", it needs to close any pending "}"
3) it needs to track if it is in a "{" block, and convert indent into ;

I think that is all possible in the lexer.

Use semantic-lex-test to show a dump of what your lexer is outputting.

Eric

On 11/03/2015 10:57 AM, Kristof Bastiaensen wrote:
> Ok, so let's say I emit a 'layout-next' token for implicit ';'.  When
> inside a layout,
> this should be part of the grammar rule, and when outside the layout, it
> should
> be ignored (like whitespace).  How do I track whether I am inside or
> outside the layout?
> Do I need to specify each rule, for example 'expression', twice, one for
> inside, one for
> outside?  Or can I set a 'variable' to track this in the parser?
>
> Regards,
> Kristof
>
> On 03-11-15 16:28, Eric Ludlam wrote:
>>
>>
>> I am replying on my phone so I'll be a bit brief.
>>
>> I was providing python as an example of whitespace handling g so you would
>> have a starting point.  At a guess, it looks like your keywords like do or
>> let also define some sort of indentation depth.  The key for your lexer is
>> to create tokens that the parser will use to reason on how it will
>> interpret the system.  Your parser would need to accept and use or ignore
>> the special tokens depending on context.  I don't think the lexer would
>> convert whitespace into braces and ; though if it isn't completely
>> deterministic from the lexical point of view.  If you can track what you
>> need through the lexer, then you could do the substitution.  The c lexer
>> handles preprocessor statements by tracking state while lexing so maybe
>> that style will work for you.
>>
>> Eric
>>
>> On Nov 3, 2015 9:48 AM, "Kristof Bastiaensen" <[email protected]
>> <mailto:[email protected]>> wrote:
>>
>>     messed the formatting up again :-( Sorry for the noise...
>>
>>     Here is a better (contrived but real) example, that demonstrates how
>>     layout depends on the parser context:
>>     f x = length (
>>            do a <- [20]
>>               return (30)) + x
>>               * 2
>>
>>     This translates to:
>>
>>     f x = length (
>>            do {a <- [20]
>>               ;return (30)}) + x
>>               * 2
>>
>>     Regards,
>>     Kristof Bastiaensen
>>
>>
>>     On 03-11-15 15:34, Kristof Bastiaensen wrote:
>>     > Actually the first problem can be solved at the lexer, by
>>     > remembering the indentation after the 'let', 'where', 'do' or 'of'
>>     > keyword.  The second problem is harder, since where to ignore
>>     > insert ';' tokens would depend on the parser context.
>>     >
>>     > The mailer messed up the formatting, so I'll try again,
>>     hopefully it is
>>     > right now:
>>     >
>>     > let a = 20
>>     >       b = 30
>>     >
>>     > and
>>     > let  a = 20
>>     >       b = 30
>>     >
>>     > are not the same thing.  The first translates to
>>     >
>>     > let {a = 20
>>     >       ;b = 30 }
>>     >
>>     > and the second to (the illegal)
>>     >
>>     > let {a = 20}
>>     >        b = 30
>>     >
>>     > And for example the following
>>     >
>>     > let a = 20
>>     >       b = 30
>>     >       in a + b
>>     >
>>     > translates to
>>     >
>>     > let {a = 20
>>     >       ;b = 30
>>     >       }in a + b
>>     >
>>     > not to
>>     >
>>     > let {a = 30
>>     >       ;b = 30
>>     >       ;in a + b}
>>     >
>>     > Which is illegal.  Where to insert ';' depends on the parser
>>     context,
>>     > it's not possible to do at the lex phase.
>>     >
>>     > Regards,
>>     > Kristof Bastiaensen
>>     >
>>     >
>>     >
>>     >
>>     >
>>     > On 03-11-15 15:21, Kristof Bastiaensen wrote:
>>     >> Hi,
>>     >>
>>     >> the problem in haskell is that simply specifying indent is not
>>     enough,
>>     >> for example:
>>     >>
>>     >> let a = 20
>>     >>         b = 30
>>     >>
>>     >> and
>>     >> let  a = 20
>>     >>         b = 30
>>     >>
>>     >> are not the same thing.  The first translates to
>>     >>
>>     >> let {a = 20
>>     >>         ;b = 30 }
>>     >>
>>     >> and the second to (the illegal)
>>     >>
>>     >> let {a = 20}
>>     >>         b = 30
>>     >>
>>     >> And for example the following
>>     >>
>>     >> let a = 20
>>     >>         b = 30
>>     >>         in a + b
>>     >>
>>     >> translates to
>>     >>
>>     >> let {a = 20
>>     >>         ;b = 30
>>     >>         }in a + b
>>     >>
>>     >> not to
>>     >>
>>     >> let {a = 30
>>     >>         ;b = 30
>>     >>         ;in a + b}
>>     >>
>>     >> Which is illegal.  Where to insert ';' depends on the parser
>>     really,
>>     >> it's hard to do at the lex phase.
>>     >>
>>     >> Regards,
>>     >> Kristof Bastiaensen
>>     >>
>>     >>
>>     >>
>>     >> On 03-11-15 13:22, Eric Ludlam wrote:
>>     >>> On 11/02/2015 07:13 AM, Kristof Bastiaensen wrote:
>>     >>>> Hi,
>>     >>>>
>>     >>>> I'd like to add support for haskell to semantic.  Haskell
>>     syntax is
>>     >>>> sensitive
>>     >>>> to indentation, and I'd like to know if it is possible to use
>>     bovine or
>>     >>>> wisent for
>>     >>>> it.
>>     >>> Yes.
>>     >>>
>>     >>>> The layout rule is explained here:
>>     >>>>
>>     https://www.haskell.org/onlinereport/haskell2010/haskellch2.html#x7-210002.7
>>     >>>>
>>     >>>>
>>     >>>> What it does, is after certain keywords, whenever a '{'
>>     character is
>>     >>>> ommited, it
>>     >>>> enters a layout context.  In the layout context it will
>>     remember the
>>     >>>> indentation
>>     >>>> of the next token.  When a line has the same indentation as
>>     this token,
>>     >>>> a ';' token is
>>     >>>> inserted before the expression.  Whenever the indentation is
>>     less than
>>     >>>> this, a '}' token
>>     >>>> is inserted, and the layout context is ended.  The layout
>>     context is
>>     >>>> also ended when a parser error would occur.
>>     >>> The python grammar in semantic/wisent/python.wy and python.el
>>     is an ok
>>     >>> example.  The premise is that when you develop the lexer, you
>>     teach it
>>     >>> to convert your indentation into lexical tokens that you then
>>     use in
>>     >>> your grammar.
>>     >>>
>>     >>> In the python.el support file, you will notice it uses
>>     >>> (current-indentation) and uses that to derive INDENT and DEDENT
>>     >>> tokens.  The grammar uses those to create INDENT and DEDENT
>>     blocks so
>>     >>> it can recurse into code bodies.  You could just as easily
>>     skip them
>>     >>> too if you prefer.
>>     >>>
>>     >>>> So I would set a state variable to all the layout context
>>     indentations,
>>     >>>> and then change the behaviour of the lexer or parser based on
>>     this
>>     >>>> state.  Is that possible?  Would it interfere with
>>     incremental parsing?
>>     >>> Definitely start with the lexer, and there is good
>>     documentation on
>>     >>> creating a lexer in the semantic language developers guide.
>>     Use the
>>     >>> lexer testing function to make sure it outputs what you
>>     expect, and
>>     >>> then you can continue on with the grammar.
>>     >>>
>>     >>> Eric
>>     >>
>>     ------------------------------------------------------------------------------
>>     >> _______________________________________________
>>     >> cedet-semantic mailing list
>>     >> [email protected]
>>     <mailto:[email protected]>
>>     >> https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>>     >
>>     >
>>     ------------------------------------------------------------------------------
>>     > _______________________________________________
>>     > cedet-semantic mailing list
>>     > [email protected]
>>     <mailto:[email protected]>
>>     > https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>>
>>
>>     ------------------------------------------------------------------------------
>>     _______________________________________________
>>     cedet-semantic mailing list
>>     [email protected]
>>     <mailto:[email protected]>
>>     https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>>
>
>
> ------------------------------------------------------------------------------
>
>
>
> _______________________________________________
> cedet-semantic mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/cedet-semantic
>

------------------------------------------------------------------------------