Data Import Policy
Andre Wiethoff <[email protected]> Mon, 11 May 2015 12:24:57 +0200
| Newsgroups | gmane.comp.audio.musicbrainz.devel |
|---|---|
| Message-ID | <[email protected]> |
Hello everybody,
its again me, with some other weird ideas ;-)
I would like to ask whether an automatic metadata collection/crawling
for insertion in the Musicbrainz DB is fine, which will be created by
web data mining?
Basically I would like to work on these two automatic data
crawlings/data minings:
1) Add links for more artists to the AMG, Amazon and BBC web pages
(which will be automatically be matched by crawling the appropriate web
pages - of course very conservatively). Also interesting would be links
to the "Musixmatch" lyrics website ( https://www.musixmatch.com ). It
seems legit, so can it be added to the lyrics site whitelist?
2) Adding a new kind of relation (which also would need to be approved
first), I would call it "similar to". Basically I would mainly
automatically add similarities between two artists, but of course also
other similarities would be possible (song similarity, etc.). As this
kind of data is highly subjective, it might be a thought whether only
data by automatic web/database data mining would be accepted as input
(and no manual input of users)... This kind of data is available on AMG,
Amazon and BBC, which could be automatically be crawled (and only two
artist IDs would be added to the database as similar). Of course lateron
similarity could also be calculated by an algorithm using some
scrobbling data.
Are the two scenarios permitted by the data import policy of Musicbrainz
(e.g. doesn't violate any copyright issues, etc.). I think the first
case shouldn't create any problems at all, as using a link all
references are given to the data source (it is just a link).
The second case is a bit more difficult and problematic, as basically
the data is created by the appropriate companies and inserted into a new
database owned by somebody different? On the other hand, only two IDs
are stored (IDs which only makes sense in Musicbrainz) - can that data
violate copyright issues? I think it is a bit similar to Google crawling
and provide the results as their own...
What do you think?
Best regards,
Andre
PS: Here a small evaluation of the URLs stored in the system for some
link targets:
URLs total: 2117159
Discogs total: 567330
Discogs Release: 212163
Discogs Artist: 200194
Discogs Master: 126747
Discogs label: 26405
Allmusic total: 55395
Allmusic Artist: 29943
Allmusic Album: 20290
Allmusic Composition: 4216
Amazon total: 183044
Amazon Product: 182683
Amazon Artist: 229
BBC total: 9805
BBC Artist: 1347
BBC Reviews: 8208
Soundcloud: 26166
Youtube total: 29747
Youtube User Channel: 15127
Youtube Video: 12723
I wonder why there are not more BBC artist links, as these links are
just http://www.bbc.co.uk/music/artists/<musicbrainz artist id>. It
should be pretty easy to add them...