Re: Next YAML: drop equality definition

Andrey Somov <[email protected]> Mon, 7 Mar 2016 16:42:16 +0100
Newsgroups gmane.text.yaml.general
Message-ID <CALbkv0dqzTq4X6w6w0YGJjTyS+FN4rkDi9teva_6PX3TMaW1jQ@mail.gmail.com>
>Oren, if you are suggesting that {a: 1, a: 2} *must* be detected at the
parser level, then I'd have to disagree from the streaming >standpoint,
since the 2 keys may be lightyears away from each other.

I completely support that {a: 1, a: 2} must NOT be detected at the parser
level.

I must admit that I do not follow the discussion here because it is too
much,
but I agree with Osamu.

Cheers,
Andrey



On Mon, Mar 7, 2016 at 4:27 PM, Ingy dot Net <[email protected]> wrote:

> Oren, Osamu,
>
> My concern with this thread is that you seem to be limiting your meaning
> of "YAML processors" to Loader/Dumper processors (processors that carry
> data from text to native programming in-memory objects, and back). ie Where
> the "Application" level uses native objects.
>
> While that is certainly the normal case, YAML conjecture should also be
> weighed against processors that stop after the parser (at the event
> stream). ie Where the application is a streaming filter/mapper. ie No graph
> construction or native objects ever happen.
>
> YAML conjecture should also be weighed against non-terminating (or
> extremely huge) YAML nodes (documents, mappings, sequences).
>
> Oren, if you are suggesting that {a: 1, a: 2} *must* be detected at the
> parser level, then I'd have to disagree from the streaming standpoint,
> since the 2 keys may be lightyears away from each other.
>
> FWIW, libyaml doesn't reason on keys and shouldn't. libyaml makes a great
> streaming parser/emitter.
>
> Key order can't really even be discussed at this level. Key order *can't* be
> changed by a parser, so of course you can make it meaningful if you decide
> to. You just have to be aware that your meaningfulness expires should your
> YAML be used in a graphing context.
>
> More below:
>
> On Sun, Mar 6, 2016 at 10:27 AM, Oren Ben-Kiki <[email protected]> wrote:
>
>> On Sun, Mar 6, 2016 at 7:00 PM, Osamu TAKEUCHI <[email protected]> wrote:
>>
>>> Oren,
>>>
>>> I expect we share the same thinking that the
>>> definition of equality belongs to the domain
>>> specific data type, not to the serialization
>>> language.
>>
>>
>> Pretty much.
>>
>>
>>> So, unless it makes the serialized
>>> documents much more readable or portable, a
>>> serialization language should not determine
>>> its own equality or identity definition.
>>>
>>
>> That's a bit "unless".
>>
>>
>>> I agree duplicate key should be detected by
>>> YAML processors because we do not want users
>>> to use duplicate keys for overwriting the
>>> values of predefined keys.
>>
>>
> There is no concept of overwriting in a streaming processor. You need a
> graph to do that.
>
>
>>
>> Yes.
>>
>>
>>> The key order in
>>> a YAML mapping should not have meaning.
>>>
>>
>> Very strong yes.
>>
>
> Depends on the layer.
>
>
>>
>>
>>> But it can be done without defining equality
>>> in YAML spec. YAML processor can use native
>>> equality evaluator of the data at its
>>> construction stage and it should do so.
>>>
>>
>> No. You assume there _is_ a construction stage. There need not be one.
>>
>
> Oren. I agree, although I think you mean that construction is skipped on
> the way to native, where I'm saying that processing might never get to that
> level.
>
>
>>
>>
>>> If the layered structure of the YAML processor
>>> do not allow it, the layered structure itself
>>> should be revised. I don't see how the layered
>>> structure is related to the current topic,
>>> though.
>>>
>>
>> It is crucial to the discussion. We all agree that key duplication
>> detection _must_ be done at the application layer, but the point is that a
>> _limited_ form of key duplication detection _may_ and _should_ be done,
>> especially in YAML processors that do not even _have_ an application layer.
>> This is because, as you put it, "it makes the serialized documents much
>> more readable or portable".
>>
>
> If the application is not a graph then duplication mustn't and likely
> can't be done.
>
>
>>
>>
>>> Similarly, I do not want to forbid PHP users
>>> to store a PHP's native key-order-aware hash
>>> into a key-order-unaware YAML mapping
>>
>>
>> The problem is, how can you tell whether this is/not safe to do? When
>> dumping such a hash table to YAML, the application needs to provide some
>> hint to the YAML processor whether this is actually safe. By default, it is
>> _not_ safe, so without an explicit hint, the YAML processor _should_ do the
>> safe thing and emit it as an !!omap.
>>
>
> PHP is not the outlier. Most JavaScript implementations have predictable
> key order. I use JS all the time to preserve key order when doing things
> like converting JSON to (block formatted) YAML.
>
>
>>
>>> Meaningfulness of the data identity should
>>> also belong to the specific data types. As
>>> shown by the previous examples, the difference
>>> in the semantics of a scalar node and that of
>>> a complex node is not always clear.
>>
>>
>> Looks pretty clear to me. Scalars are "values". They have _no_ identity,
>> they have _only_ content. Complex nodes have identity, and as you pointed
>> out, this means their actual content may be irrelevant (for comparison).
>> The current spec gets that last point wrong.
>>
>> I think the core issue here is identity of scalars. You seem to assume
>> that a YAML processor _must_ preserve the identity of scalars. That is, it
>> _must not_, for example, use interned strings for keys. The current spec
>> says the opposite. A YAML processor _need not_ preserve scalar identity and
>> it _may_ use interned strings and other similar tricks. It is definitely
>> not required to keep the identity of, say, integer scalars!
>>
>> ... without declaring
>>> possible non-preservation of identity for
>>> scalars, nobody will think a data with an
>>> identity-based equality evaluation must be
>>> stored as a collection node and must not as a
>>> scalar node. It brings some surprise to users.
>>>
>>
>> Really? People would be very surprised to hear that { a: 1, a: 2 } is
>> actually OK because some application somewhere _may_ decide it wants scalar
>> string keys to use "identity-based equality".
>>
>> So *No*.
>>
>> For increased readability and portability, the above _must_ be allowed to
>> be flagged as a duplicate by a YAML processor regardless of what the
>> application is. And _should_ be flagged so by "well behaved" processors.
>> Even if they do _not_ have an application layer.
>>
>
> Whoops, I misread your sentence. "_must_ be allowed to be flagged". Yes I
> completely agree with that statement. *If* a parser wants to detect
> duplicate content/tag scalar keys it must be *allowed* to. I first read
> this to say that parsers must. Since you and I have considered streaming
> from the start, I was a bit surprised that you would say that. :)
>
> My other points stand.
>
> ...
>
>
>
>>
>>> Such restriction will improve YAML's readability
>>> and portability very little if any.
>>
>>
>> We'll have to agree to disagree, I'm afraid...
>>
>>
>>> Actually, I
>>> believe the restriction is currently not known
>>> widely and very few libraries and applications
>>
>> have ever utilized it.
>>
>>
>> A pity.
>>
>>
>>> I don't think many
>>> existing YAML document loose its meaning if we
>>> drop the restriction.
>>
>>
>> No valid YAML documents will, that's for sure ;-) But that's besides the
>> point.
>>
>>
>>> So, let's make the spec simpler by dropping the
>>> definition of YAML's own equality and identity
>>> preservation.
>>>
>>
>> There's no such thing as not addressing the issue of identity and
>> equality in the spec. Either you _require_ a YAML processor to preserve the
>> identity of scalars (including, horribly, simple integers), or you do not.
>> Either way it needs to be stated in the spec.
>>
>> We chose to say a YAML processer _need not_ preserve the identity of
>> scalars. Given this, then an application _must not_ use scalar identity for
>> equality comparisons. Given this, then _regardless_ of the application's
>> definition of equality, we can predict with 100% certainty that { a: 1, a:
>> 2 } contains a duplicate key.
>>
>> What I want to say in the spec is:
>>> A well-behaved processor _should_ detect a
>>> duplicate key and flag it as an error if it
>>> can correctly evaluate equality of nodes.
>>>
>>
>> So far so good.
>>
>>
>>> It _must_ aware that a data with some specific
>>> tags may have some custom comparison algorithms,
>>> including the one based on the data identity.
>>>
>>
>> Yes, the current spec gets the identity point wrong.
>>
>>
>>> Namely, two YAML nodes of same values and same
>>> tags can be evaluated to be unequal by an
>>> identity-based evaluator,
>>
>>
>> _Only_ if these are complex nodes.
>>
>>
>>> while two YAML nodes
>>> of different values and even different tags can
>>> be evaluated to be equal by some specific
>>> evaluators. Note that javascript do not
>>> natively distinguish an integer 0x01 with a
>>> sequence [1] as mapping keys.
>>>
>>
>> You keep conflating false positives with false negatives. False negatives
>> are _fine_. It is OK for the processor to miss some cases of key
>> duplication. In fact it is expected. The application is the final arbiter
>> of key equality. You can keep on piling as many examples of "the processor
>> can't detect keys in <some example> as duplicated" as you want. OF COURSE
>> there are such cases.
>>
>> But this does not mean in any way shape or form that we allow false
>> positives. A YAML processor must not ever complain about key duplication
>> when such a duplication does not exist. Now, in JavaScript, PHP, Perl,
>> Ruby, Python, C++, and any other valid YAML system, { a: 1, a: 2 } _does_
>> have a duplicate key. So a processor _is_ allowed and _should_ complain
>> about this, _regardless_ of the application-defined equality operator.
>>
>> It is also warned that tags of nodes can be
>>> implicitly specified by the path of the node
>>> from the root. So, a schema-blind YAML processor
>>> can never know how to resolve a tag for any
>>> tag-unspecified node.
>>
>>
>> The path to both the "a" keys in { a: 1, a: 2 } is, by definition, the
>> same (the path to all keys in the same mapping is, by definition,
>> identical). So whatever tag is assigned to one of them, by definition, the
>> same tag must be assigned to the other as well. The application _can't_ use
>> different tags to distinguish between them. It _can't_ use their identity
>> to distinguish between them because the YAML processor need not give them
>> different identities. It _can't_ use their content to distinguish between
>> them because they have the same content. So, the application _must_
>> consider them equal - there's just no other possible choice.
>>
>> So we _can_ complain about them being equal at an earlier processing
>> stage. We do not _require_ a YAML processor to do so, but we _allow_ and
>> _encourage_ it to do so.
>>
>>
>>> A well-behaved YAML
>>> processor _must_ be schema aware,
>>
>>
>> Now this is just plain wrong. YamlReference is a YAML processor. It
>> implements the parsing stage. It has no clue whatsoever what schema is
>> used. Schema-blind YAML processing is, for me, an important use case.
>>
>> And if every possible schema in the universe _must_ decree that two keys
>> are equal, then we don't need to know the _specific_ schema, because
>> whatever it is, it will also _have_ to declare them equal.
>>
>> Oren.
>>
>>
>> ------------------------------------------------------------------------------
>>
>> _______________________________________________
>> Yaml-core mailing list
>> [email protected]
>> https://lists.sourceforge.net/lists/listinfo/yaml-core
>>
>>
>
>
> ------------------------------------------------------------------------------
> Transform Data into Opportunity.
> Accelerate data analysis in your applications with
> Intel Data Analytics Acceleration Library.
> Click to learn more.
> http://makebettercode.com/inteldaal-eval
> _______________________________________________
> Yaml-core mailing list
> [email protected]
> https://lists.sourceforge.net/lists/listinfo/yaml-core
>
>

------------------------------------------------------------------------------
Transform Data into Opportunity.
Accelerate data analysis in your applications with
Intel Data Analytics Acceleration Library.
Click to learn more.
http://makebettercode.com/inteldaal-eval

_______________________________________________
Yaml-core mailing list
[email protected]
https://lists.sourceforge.net/lists/listinfo/yaml-core