Re: Next YAML: drop equality definition

Ingy dot Net <[email protected]> Mon, 7 Mar 2016 07:27:20 -0800
Newsgroups gmane.text.yaml.general
Message-ID <CAHJtQJ5OLVv93o5vsBO6zvqWXFw7TJPZUvNsi_GiAKV=u-qaqQ@mail.gmail.com>
Oren, Osamu,

My concern with this thread is that you seem to be limiting your meaning of
"YAML processors" to Loader/Dumper processors (processors that carry data
from text to native programming in-memory objects, and back). ie Where the
"Application" level uses native objects.

While that is certainly the normal case, YAML conjecture should also be
weighed against processors that stop after the parser (at the event
stream). ie Where the application is a streaming filter/mapper. ie No graph
construction or native objects ever happen.

YAML conjecture should also be weighed against non-terminating (or
extremely huge) YAML nodes (documents, mappings, sequences).

Oren, if you are suggesting that {a: 1, a: 2} *must* be detected at the
parser level, then I'd have to disagree from the streaming standpoint,
since the 2 keys may be lightyears away from each other.

FWIW, libyaml doesn't reason on keys and shouldn't. libyaml makes a great
streaming parser/emitter.

Key order can't really even be discussed at this level. Key order *can't* be
changed by a parser, so of course you can make it meaningful if you decide
to. You just have to be aware that your meaningfulness expires should your
YAML be used in a graphing context.

More below:

On Sun, Mar 6, 2016 at 10:27 AM, Oren Ben-Kiki <[email protected]> wrote:

> On Sun, Mar 6, 2016 at 7:00 PM, Osamu TAKEUCHI <[email protected]> wrote:
>
>> Oren,
>>
>> I expect we share the same thinking that the
>> definition of equality belongs to the domain
>> specific data type, not to the serialization
>> language.
>
>
> Pretty much.
>
>
>> So, unless it makes the serialized
>> documents much more readable or portable, a
>> serialization language should not determine
>> its own equality or identity definition.
>>
>
> That's a bit "unless".
>
>
>> I agree duplicate key should be detected by
>> YAML processors because we do not want users
>> to use duplicate keys for overwriting the
>> values of predefined keys.
>
>
There is no concept of overwriting in a streaming processor. You need a
graph to do that.


>
> Yes.
>
>
>> The key order in
>> a YAML mapping should not have meaning.
>>
>
> Very strong yes.
>

Depends on the layer.


>
>
>> But it can be done without defining equality
>> in YAML spec. YAML processor can use native
>> equality evaluator of the data at its
>> construction stage and it should do so.
>>
>
> No. You assume there _is_ a construction stage. There need not be one.
>

Oren. I agree, although I think you mean that construction is skipped on
the way to native, where I'm saying that processing might never get to that
level.


>
>
>> If the layered structure of the YAML processor
>> do not allow it, the layered structure itself
>> should be revised. I don't see how the layered
>> structure is related to the current topic,
>> though.
>>
>
> It is crucial to the discussion. We all agree that key duplication
> detection _must_ be done at the application layer, but the point is that a
> _limited_ form of key duplication detection _may_ and _should_ be done,
> especially in YAML processors that do not even _have_ an application layer.
> This is because, as you put it, "it makes the serialized documents much
> more readable or portable".
>

If the application is not a graph then duplication mustn't and likely can't
be done.


>
>
>> Similarly, I do not want to forbid PHP users
>> to store a PHP's native key-order-aware hash
>> into a key-order-unaware YAML mapping
>
>
> The problem is, how can you tell whether this is/not safe to do? When
> dumping such a hash table to YAML, the application needs to provide some
> hint to the YAML processor whether this is actually safe. By default, it is
> _not_ safe, so without an explicit hint, the YAML processor _should_ do the
> safe thing and emit it as an !!omap.
>

PHP is not the outlier. Most JavaScript implementations have predictable
key order. I use JS all the time to preserve key order when doing things
like converting JSON to (block formatted) YAML.


>
>> Meaningfulness of the data identity should
>> also belong to the specific data types. As
>> shown by the previous examples, the difference
>> in the semantics of a scalar node and that of
>> a complex node is not always clear.
>
>
> Looks pretty clear to me. Scalars are "values". They have _no_ identity,
> they have _only_ content. Complex nodes have identity, and as you pointed
> out, this means their actual content may be irrelevant (for comparison).
> The current spec gets that last point wrong.
>
> I think the core issue here is identity of scalars. You seem to assume
> that a YAML processor _must_ preserve the identity of scalars. That is, it
> _must not_, for example, use interned strings for keys. The current spec
> says the opposite. A YAML processor _need not_ preserve scalar identity and
> it _may_ use interned strings and other similar tricks. It is definitely
> not required to keep the identity of, say, integer scalars!
>
> ... without declaring
>> possible non-preservation of identity for
>> scalars, nobody will think a data with an
>> identity-based equality evaluation must be
>> stored as a collection node and must not as a
>> scalar node. It brings some surprise to users.
>>
>
> Really? People would be very surprised to hear that { a: 1, a: 2 } is
> actually OK because some application somewhere _may_ decide it wants scalar
> string keys to use "identity-based equality".
>
> So *No*.
>
> For increased readability and portability, the above _must_ be allowed to
> be flagged as a duplicate by a YAML processor regardless of what the
> application is. And _should_ be flagged so by "well behaved" processors.
> Even if they do _not_ have an application layer.
>

Whoops, I misread your sentence. "_must_ be allowed to be flagged". Yes I
completely agree with that statement. *If* a parser wants to detect
duplicate content/tag scalar keys it must be *allowed* to. I first read
this to say that parsers must. Since you and I have considered streaming
from the start, I was a bit surprised that you would say that. :)

My other points stand.

...



>
>> Such restriction will improve YAML's readability
>> and portability very little if any.
>
>
> We'll have to agree to disagree, I'm afraid...
>
>
>> Actually, I
>> believe the restriction is currently not known
>> widely and very few libraries and applications
>
> have ever utilized it.
>
>
> A pity.
>
>
>> I don't think many
>> existing YAML document loose its meaning if we
>> drop the restriction.
>
>
> No valid YAML documents will, that's for sure ;-) But that's besides the
> point.
>
>
>> So, let's make the spec simpler by dropping the
>> definition of YAML's own equality and identity
>> preservation.
>>
>
> There's no such thing as not addressing the issue of identity and equality
> in the spec. Either you _require_ a YAML processor to preserve the identity
> of scalars (including, horribly, simple integers), or you do not. Either
> way it needs to be stated in the spec.
>
> We chose to say a YAML processer _need not_ preserve the identity of
> scalars. Given this, then an application _must not_ use scalar identity for
> equality comparisons. Given this, then _regardless_ of the application's
> definition of equality, we can predict with 100% certainty that { a: 1, a:
> 2 } contains a duplicate key.
>
> What I want to say in the spec is:
>> A well-behaved processor _should_ detect a
>> duplicate key and flag it as an error if it
>> can correctly evaluate equality of nodes.
>>
>
> So far so good.
>
>
>> It _must_ aware that a data with some specific
>> tags may have some custom comparison algorithms,
>> including the one based on the data identity.
>>
>
> Yes, the current spec gets the identity point wrong.
>
>
>> Namely, two YAML nodes of same values and same
>> tags can be evaluated to be unequal by an
>> identity-based evaluator,
>
>
> _Only_ if these are complex nodes.
>
>
>> while two YAML nodes
>> of different values and even different tags can
>> be evaluated to be equal by some specific
>> evaluators. Note that javascript do not
>> natively distinguish an integer 0x01 with a
>> sequence [1] as mapping keys.
>>
>
> You keep conflating false positives with false negatives. False negatives
> are _fine_. It is OK for the processor to miss some cases of key
> duplication. In fact it is expected. The application is the final arbiter
> of key equality. You can keep on piling as many examples of "the processor
> can't detect keys in <some example> as duplicated" as you want. OF COURSE
> there are such cases.
>
> But this does not mean in any way shape or form that we allow false
> positives. A YAML processor must not ever complain about key duplication
> when such a duplication does not exist. Now, in JavaScript, PHP, Perl,
> Ruby, Python, C++, and any other valid YAML system, { a: 1, a: 2 } _does_
> have a duplicate key. So a processor _is_ allowed and _should_ complain
> about this, _regardless_ of the application-defined equality operator.
>
> It is also warned that tags of nodes can be
>> implicitly specified by the path of the node
>> from the root. So, a schema-blind YAML processor
>> can never know how to resolve a tag for any
>> tag-unspecified node.
>
>
> The path to both the "a" keys in { a: 1, a: 2 } is, by definition, the
> same (the path to all keys in the same mapping is, by definition,
> identical). So whatever tag is assigned to one of them, by definition, the
> same tag must be assigned to the other as well. The application _can't_ use
> different tags to distinguish between them. It _can't_ use their identity
> to distinguish between them because the YAML processor need not give them
> different identities. It _can't_ use their content to distinguish between
> them because they have the same content. So, the application _must_
> consider them equal - there's just no other possible choice.
>
> So we _can_ complain about them being equal at an earlier processing
> stage. We do not _require_ a YAML processor to do so, but we _allow_ and
> _encourage_ it to do so.
>
>
>> A well-behaved YAML
>> processor _must_ be schema aware,
>
>
> Now this is just plain wrong. YamlReference is a YAML processor. It
> implements the parsing stage. It has no clue whatsoever what schema is
> used. Schema-blind YAML processing is, for me, an important use case.
>
> And if every possible schema in the universe _must_ decree that two keys
> are equal, then we don't need to know the _specific_ schema, because
> whatever it is, it will also _have_ to declare them equal.
>
> Oren.
>
>
> ------------------------------------------------------------------------------
>
> _______________________________________________
> Yaml-core mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/yaml-core
>
>

------------------------------------------------------------------------------
Transform Data into Opportunity.
Accelerate data analysis in your applications with
Intel Data Analytics Acceleration Library.
Click to learn more.
http://makebettercode.com/inteldaal-eval

_______________________________________________
Yaml-core mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/yaml-core