Data Import Policy

Andre Wiethoff <[email protected]> Mon, 11 May 2015 12:24:57 +0200
Newsgroups gmane.comp.audio.musicbrainz.devel
Message-ID <[email protected]>
Hello everybody,

its again me, with some other weird ideas ;-)
I would like to ask whether an automatic metadata collection/crawling 
for insertion in the Musicbrainz DB is fine, which will be created by 
web data mining?

Basically I would like to work on these two automatic data 
crawlings/data minings:

1) Add links for more artists to the AMG, Amazon and BBC web pages 
(which will be automatically be matched by crawling the appropriate web 
pages - of course very conservatively). Also interesting would be links 
to the "Musixmatch" lyrics website ( https://www.musixmatch.com ). It 
seems legit, so can it be added to the lyrics site whitelist?

2) Adding a new kind of relation (which also would need to be approved 
first), I would call it "similar to". Basically I would mainly 
automatically add similarities between two artists, but of course also 
other similarities would be possible (song similarity, etc.). As this 
kind of data is highly subjective, it might be a thought whether only 
data by automatic web/database data mining would be accepted as input 
(and no manual input of users)... This kind of data is available on AMG, 
Amazon and BBC, which could be automatically be crawled (and only two 
artist IDs would be added to the database as similar). Of course lateron 
similarity could also be calculated by an algorithm using some 
scrobbling data.

Are the two scenarios permitted by the data import policy of Musicbrainz 
(e.g. doesn't violate any copyright issues, etc.). I think the first 
case shouldn't create any problems at all, as using a link all 
references are given to the data source (it is just a link).
The second case is a bit more difficult and problematic, as basically 
the data is created by the appropriate companies and inserted into a new 
database owned by somebody different? On the other hand, only two IDs 
are stored (IDs which only makes sense in Musicbrainz) - can that data 
violate copyright issues? I think it is a bit similar to Google crawling 
and provide the results as their own...

What do you think?

Best regards,

Andre

PS: Here a small evaluation of the URLs stored in the system for some 
link targets:

URLs total: 2117159
    Discogs total: 567330
       Discogs Release: 212163
       Discogs Artist: 200194
       Discogs Master: 126747
       Discogs label: 26405
    Allmusic total: 55395
       Allmusic Artist: 29943
       Allmusic Album: 20290
       Allmusic Composition: 4216
    Amazon total: 183044
       Amazon Product: 182683
       Amazon Artist: 229
    BBC total: 9805
       BBC Artist: 1347
       BBC Reviews: 8208

    Soundcloud: 26166
    Youtube total: 29747
       Youtube User Channel: 15127
       Youtube Video: 12723

I wonder why there are not more BBC artist links, as these links are 
just http://www.bbc.co.uk/music/artists/<musicbrainz artist id>. It 
should be pretty easy to add them...