Re: GSoC 2016 Letor dataset discussion

Parth Gupta <[email protected]>
Newsgroups gmane.comp.search.xapian.devel
Message-ID <CAMsucs3W_=jRW1Qgc1KyStrLo1HL64La_v2kD65mOVn3wJrSKA@mail.gmail.com>
I used a subset of INEX 2009 with around 2M documents (some details here:
https://trac.xapian.org/wiki/GSoC2011/LTR/Notes#IREvaluationofLetorrankingscheme)
and it worked fine. If you have access to it, should work for most of our
purposes.

As the INEX documents have rich xml meta-data, letor can benefit in terms
of fields (title, body etc.)

For unit-testing, as James mentions, go with automated tests in a
controlled environment. Use INEX data-set for explicit evaluation and see
if everything works without breaking at large scale.

Cheers
Parth

On Sat, May 14, 2016 at 9:57 PM, James Aylett <[email protected]>
wrote:

> On Sat, May 14, 2016 at 04:51:57PM +0530, Ayush Tomar wrote:
>
> > I wanted to decide the dataset that should be used for Letor
> stabilisation
> > project.
>
> Is this for evaluating the various letor approaches? For unit tests
> you'll need to generate your own test data (partly so you can control
> it better to do validation properly, but also because the licenses
> almost never work).
>
> Parth should be able to advise on suitable datasets for evaluating
> letor.
>
> J
>
> --
>   James Aylett, occasional trouble-maker
>   xapian.org
>
>
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.