Re: Data Import Policy
Andre Wiethoff <[email protected]> Tue, 12 May 2015 16:00:13 +0200
| Newsgroups | gmane.comp.audio.musicbrainz.devel |
|---|---|
| Message-ID | <[email protected]> |
Hello Frederik, thanks for your thoughts! > > This should not be a relationship like we currently have relationships. > This is an entirely subjective piece of information and some people will > consider two artists similar while others will not. Yes, therefore I did propose to not open it for public editing, but = having it "computer generated" only (by whatever means, be it web = crawling or clustering on large data sets). > There are already > projects (though I forget which, sorry) that group/cluster entities > based on relationships (which IIRC wasn't completely off), so it is > possible to do something like it with the data in MB already. I wonder which relationships have been used to group/cluster the = entities? With the existing Metabrainz data, I can think only of = scrobbling data to generate such kind of information... I found this paper, but they used Musicbrainz only for the basic = metadata retrieval and Audioscrobbler (last.fm) for the similarity = clustering... http://www.sfu.ca/~shaw/papers/musicianMap-VDA09.pdf > There's > also AcousticBrainz which can be used to cluster entities based on the > acoustic properties of their recordings. I don't think that clustering regarding the acoustic properties will = bring any good results for now, I guess this will still take ten years = until there exist something that produces results matchable to a human = expert (or even advanced amateur)... > If we get scrobbling hooked in > to our end at some point, that'd obviously be another usable source for > this, but it isn't required to get something like this going. But in the end you agree that the result of such a web crawl/clustering = algorithm/whatever should be stored in the database as final result (for = speedier access of the results) - if implemented at all? But perhaps we = should discuss at first whether the new data would be beneficial for the = users (or the database)... I thought that relationships would have been the best place to put them, = as in fact it is a relation between e.g. two artists (even though the = definition of similarity would be depend on the algorithm or the page = that is crawled). E.g. Amazon will most probably use the "customers that = buy stuff from this artist also buyed stuff from these other artists" = similarity measurement. I am unsure which measurements are used by AMG = and BBC, but most probably also some kind of clustering algorithm... Please see the similarity results of the pages for the artist "Herbert = Gr=F6nemeyer": http://www.allmusic.com/artist/herbert-gr%C3%B6nemeyer-mn0000956217/related http://www.amazon.de/Herbert-Groenemeyer/e/B000APL43M http://www.bbc.co.uk/music/artists/456eabce-d1dd-4481-a206-36ab4f2eaeb8#more >> I think it is a bit similar to Google crawling >> and provide the results as their own... > Google does pay at least some of their data sources. I know, for one, > that Google is MetaBrainz' biggest "customer" by far (in terms of how > much money they put in the project). > I also see this a bit controversial, as even two IDs could be = intellectual property... I think that it is the biggest question of whether to allow web crawling = for this purpose at all. Does anybody else have some insights on this? Thank your in forward for your answers! Best regards, Andre