Re: StringBuilder Extension: IsQuotedBy

Peter Obiefuna <[email protected]> Fri, 15 Feb 2008 10:55:11 -0700
Newsgroups gmane.comp.windows.devel.dotnet.clr
Message-ID <[email protected]>
This is an illuminating generalization, Frans. However, my response is in
the context of whether to whip up a home-grown DFA to check if a string
inside a .NET StringBuilder contains a quote at it's end or not.
P

--------------------------------------------------
From: "Frans Bouma" <[email protected]>
Sent: Friday, February 15, 2008 10:39 AM
To: <[email protected]>
Subject: Re: [DOTNET-CLR] StringBuilder Extension: IsQuotedBy

>> Peter Obiefuna <[email protected]> wrote:
>> > If it keeps hitting a
>> > matched state until the input buffer is finished, then, the string
>> > 'qualifies'. That, in my mind, is the difference between  a state
>> > engine
> and
>> > a collation engine like strcomp (never mind that you can illustrate FSM
>> > graphically by pointing to a final dot on paper). But I expect a Regex
>> > implementation to create a unique FSM from every input signature.
>>
>> I don't understand your last sentence. The FSM is constructed for the
>> regex, not for the sentence to be matched.
>>
>> Also, I don't see how this is relevant to why one wouldn't "still"
>> hand-write a DFA.
>
>        I think it depends on your language. Take for example UBB syntaxis
> vs.
> a programming language. The first has tokens which are surrounded by
> markers.
> The second doesn't. THe first has whitespace + text outside the marked
> areas
> which is the same: ignore and see it as one token, the second doesn't, it
> has
> to consider every input character.
>
>        The first can be done with a regexp tokenizer pretty easily and
> very
> quickly: one can find the tokens in the text without a lot of effort, just
> define the regexp per token. (an open source example I wrote some time ago
> is
> available here: http://www.llblgen.com/hnd) The second can't use that
> setup
> because every non-whitespace is input to tokenize. Using a regexp
> tokenizer
> would be too slow (as it's too inefficient). THe first is efficient as it
> only
> has to consider the marked areas, the rest is to ignore.
>
>        So if you're writing a parser for a language where every
> non-whitespace is a token, you indeed need your own statemachine to
> tokenize
> the input.
>
>        I think both have a point: regexp's already define NFA's interally
> for
> their expression, why not utilize those? Of course, if your language
> doesn't
> fit that setup, you need your own. (Aho Sethi Ullman to the rescue ;))
>
>                FB
>
> ===================================
> This list is hosted by DevelopMentor®  http://www.develop.com
>
> View archives and manage your subscription(s) at
> http://discuss.develop.com
>

===================================
This list is hosted by DevelopMentor®  http://www.develop.com

View archives and manage your subscription(s) at http://discuss.develop.com