Re: YAML Multi-Layer Spec
Andrey Somov <[email protected]> Sun, 6 Mar 2016 19:22:13 +0100
| Newsgroups | gmane.text.yaml.general |
|---|---|
| Message-ID | <CALbkv0e4UMnZT7o8ndiR-DoQEvGmpLsa5GHCYrH_4z-R5MOsSw@mail.gmail.com> |
It is not only PyYAML. YAML 1.2: At the end of a document <http://www.yaml.org/spec/1.2/spec.html#document//>, a *document end marker* line is used to signal the parser <http://www.yaml.org/spec/1.2/spec.html#parse//> to begin scanning for directives <http://www.yaml.org/spec/1.2/spec.html#directive//> again. "Scanning for directives" sounds natural, "lexing for directives" does not. If we need to choose form lex, tokenize and scan then my preference is: 1. scan 2. tokenize 3. lex The specification does not use term "lexeme", but uses "token". Cheers, Andrey On Sun, Mar 6, 2016 at 5:14 PM, Ingy dot Net <[email protected]> wrote: > Most parsers in the real world use the term lexer. PyYaml is the first > place I saw the term scanner used. > > On Sat, Mar 5, 2016 at 2:22 AM, Andrey Somov <[email protected]> > wrote: > >> Minor proposal: >> let us use 'scan' instead of 'lex' >> >> 1. I could not find the word 'lex', but could find 'scan' >> 2. scan sounds better >> 3. scanner is already used in PyYAML and SnakeYAML >> >> Cheers, >> Andrey >> >> >> >> On Fri, Mar 4, 2016 at 9:40 PM, Ingy dot Net <[email protected]> wrote: >> >>> Since this thread seems to be winding down, I've changed the subject. >>> About a year ago I was driving home from a conference where I gave a talk >>> on YAML. I had a sort of epiphany where I realized the source of a lot of >>> the contentions we had (and still have) around things like duplicate keys, >>> key order, equality and such. >>> >>> YAML processing is stack based: See the diagram in >>> https://github.com/ingydotnet/yaml-dev-kit/blob/guide/yaml-loader-guide.pod#readme >>> >>> A Load operation is: >>> (source)->read->(chars)->lex->(tokens)->parse->(events)->compose->(graph)->construct->(native) >>> >>> The YAML Spec talks about all of these things, but primarily it defines >>> precisely a token grammar. >>> >>> My big realization was that the spec needs to "spec" each layer and >>> transform. They are each distinct, and need to be defined and reasoned >>> about distinctly. >>> >>> For instance, it is not good enough to say "On input, a YAML processor >>> must support the UTF-8 and UTF-16 character encodings. For JSON >>> compatibility, the UTF-32 encodings must also be supported." >>> >>> You would want to state this in the "Reader" section. >>> >>> Statements about Key Order and Duplicate Keys can't be made about "YAML" >>> but about YAML processing at the graph layer. At the parser level there can >>> be no concept of key order or duplication. In a streaming processor, it is >>> almost certain that key order will be preserved. >>> >>> ... >>> >>> I don't think we need to be alarmed in any way. I think the spec is fine >>> for what it is. I am hoping to capture all of this in a more accessible >>> form in the dev kit <https://github.com/ingydotnet/yaml-dev-kit>. Also >>> with tests that force the issues. I think that will be much easier to >>> reason about than the spec. >>> >>> On Fri, Mar 4, 2016 at 9:04 AM, Oren Ben-Kiki <[email protected]> wrote: >>> >>>> >>>> >>>> On Fri, Mar 4, 2016 at 7:01 PM, Osamu TAKEUCHI <[email protected]> wrote: >>>> >>>>> Let me confirm one point. Does this allow identity-based >>>>> comparison of mapping nodes with some specific tags? >>>> >>>> >>>> Not sure what you mean. In YAML each mapping node has its own identity >>>> which is different from all other nodes. The only way for two mapping nodes >>>> to have the same identity is via an alias (anchor and reference). >>>> >>>> >>>>> > But that issue doesn't seem to be the one raised here; >>>>> >the problem seems to be with the "early duplicates detection". >>>>> >Which, again, is explicitly not required by the spec. >>>>> >So I'm not certain what the actual problem is with the current >>>>> >rules. Is this a case of the perceived rules being different >>>>> >from the actual rules in the spec? >>>>> >>>>> If you don't see any problem to relax the equality detection >>>>> as above, neither do I. >>>>> >>>> >>>> There's no need to "relax" it, it is "relaxed" already. >>>> >>>> Oren. >>>> >>>> >>>> ------------------------------------------------------------------------------ >>>> >>>> _______________________________________________ >>>> Yaml-core mailing list >>>> [email protected] >>>> https://lists.sourceforge.net/lists/listinfo/yaml-core >>>> >>>> >>> >>> >>> ------------------------------------------------------------------------------ >>> >>> _______________________________________________ >>> Yaml-core mailing list >>> [email protected] >>> https://lists.sourceforge.net/lists/listinfo/yaml-core >>> >>> >> >> >> ------------------------------------------------------------------------------ >> >> _______________________________________________ >> Yaml-core mailing list >> [email protected] >> https://lists.sourceforge.net/lists/listinfo/yaml-core >> >> > ------------------------------------------------------------------------------ _______________________________________________ Yaml-core mailing list [email protected] https://lists.sourceforge.net/lists/listinfo/yaml-core