Re: Data Import Policy

Andre Wiethoff <[email protected]> Tue, 12 May 2015 16:00:13 +0200
Newsgroups gmane.comp.audio.musicbrainz.devel
Message-ID <[email protected]>
Hello Frederik,

thanks for your thoughts!
>
> This should not be a relationship like we currently have relationships.
> This is an entirely subjective piece of information and some people will
> consider two artists similar while others will not.
Yes, therefore I did propose to not open it for public editing, but =

having it "computer generated" only (by whatever means, be it web =

crawling or clustering on large data sets).
> There are already
> projects (though I forget which, sorry) that group/cluster entities
> based on relationships (which IIRC wasn't completely off), so it is
> possible to do something like it with the data in MB already.
I wonder which relationships have been used to group/cluster the =

entities? With the existing Metabrainz data, I can think only of =

scrobbling data to generate such kind of information...
I found this paper, but they used Musicbrainz only for the basic =

metadata retrieval and Audioscrobbler (last.fm) for the similarity =

clustering...
http://www.sfu.ca/~shaw/papers/musicianMap-VDA09.pdf
> There's
> also AcousticBrainz which can be used to cluster entities based on the
> acoustic properties of their recordings.
I don't think that clustering regarding the acoustic properties will =

bring any good results for now, I guess this will still take ten years =

until there exist something that produces results matchable to a human =

expert (or even advanced amateur)...
> If we get scrobbling hooked in
> to our end at some point, that'd obviously be another usable source for
> this, but it isn't required to get something like this going.
But in the end you agree that the result of such a web crawl/clustering =

algorithm/whatever should be stored in the database as final result (for =

speedier access of the results) - if implemented at all? But perhaps we =

should discuss at first whether the new data would be beneficial for the =

users (or the database)...

I thought that relationships would have been the best place to put them, =

as in fact it is a relation between e.g. two artists (even though the =

definition of similarity would be depend on the algorithm or the page =

that is crawled). E.g. Amazon will most probably use the "customers that =

buy stuff from this artist also buyed stuff from these other artists" =

similarity measurement. I am unsure which measurements are used by AMG =

and BBC, but most probably also some kind of clustering algorithm...

Please see the similarity results of the pages for the artist "Herbert =

Gr=F6nemeyer":
http://www.allmusic.com/artist/herbert-gr%C3%B6nemeyer-mn0000956217/related
http://www.amazon.de/Herbert-Groenemeyer/e/B000APL43M
http://www.bbc.co.uk/music/artists/456eabce-d1dd-4481-a206-36ab4f2eaeb8#more
>> I think it is a bit similar to Google crawling
>> and provide the results as their own...
> Google does pay at least some of their data sources. I know, for one,
> that Google is MetaBrainz' biggest "customer" by far (in terms of how
> much money they put in the project).
>
I also see this a bit controversial, as even two IDs could be =

intellectual property...
I think that it is the biggest question of whether to allow web crawling =

for this purpose at all.
Does anybody else have some insights on this?

Thank your in forward for your answers!

Best regards,

Andre