GSoC 2016 Letor dataset discussion

Ayush Tomar <[email protected]>
Newsgroups gmane.comp.search.xapian.devel
Message-ID <CAPx_PWmuPL--OohBwTuHSUnFoFh+0dDNUA2WTyAHzt2zrT+RhQ@mail.gmail.com>
Hello,

I wanted to decide the dataset that should be used for Letor stabilisation
project.

I think 2009 INEX Wikipedia Collection
<http://www.mpi-inf.mpg.de/departments/databases-and-information-systems/software/inex/>
should work fine. It's a collection of 2,666,190 XML articles, 115 topics
<http://inex.mmci.uni-saarland.de/protected/adhoc/2009-topics.zip>, 50,275
qrel <http://inex.mmci.uni-saarland.de/protected/adhoc/2009-inex_eval.zip>
labels and has an uncompressed size of 50.75 gb (5.52 GB compressed).

Another similar alternative is 2013 INEX Wikipedia LOD Collection
<http://inex-lod.mpi-inf.mpg.de/2013/>. It's a collection of 12,216,109 XML
articles, 144 topics
<http://inex.mmci.uni-saarland.de/protected/dc/2013-ld-adhoc-topics.xml>,
14,400
qrel <http://inex.mmci.uni-saarland.de/protected/dc/2013-ld-adhoc-qrels.zip>
labels. It has a compressed size of 11.12 GB. INEX 2009 Collection is a
subset of it.

If there are any recent/better datasets that can be used, please let me
know.

Thanks,
Ayush
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.