Re: GSoC 2014: Solr search

Wieland Hoffmann <[email protected]>
Newsgroups gmane.comp.audio.musicbrainz.devel
Message-ID <[email protected]>
Hallo Paul Taylor:
> On 20/03/2014 18:51, Wieland Hoffmann wrote:
> >Hallo Paul Taylor:
> >>On 19/03/2014 15:28, Wieland Hoffmann wrote:
> >>>Hallo Paul Taylor:
> >>>>Nor would I but when you say 'Solr comes with tokenizers and filters out
> >>>>of the box' it comes across to me that you think that Solr is providing
> >>>>some improvement for tokenixzers that we dont have with the existing
> >>>>system, and I was guessing at what you thought the advantage was - Im
> >>>>still at a loss to see what it is.
> >>>As I have mentioned in my proposal: I do not intend to spend much time
> >>>on perfecting the analysis process and have not invested much time in
> >>>reading up on what the current search server does and what Solr can do
> >>>by itself. I did read the reference guide, though, and the list of
> >>>provided tokenizers and filters did look good enough for most (not all)
> >>>of what's currently being done (at least from what I understood from the
> >>>documentation - this didn't work so well for the MusicBrainzTokenizer,
> >>>so it's certainly possible that my impression is completely wrong).
> >>>
> >>>The advantage of using tokenizers and filters provided by Solr itself is
> >>>not that they're going to make everything work better (maybe they will,
> >>>maybe they won't, I don't know), it's that we're reducing the amount of
> >>>code we have to maintain ourselves.
> >>>
> >>I think you are still missing my point, AFAIK Solr doesnt provide any
> >>additional filters or tokenizers its just allowing you to use the ones
> >>already provided by Lucene. So we can already use these tokenizers, we
> >>already do use some of them we only write our own tokenizers/analysers
> >>when one is not readily available within lucene, and often this just
> >>involves subclassing and existing class and modifying the required
> >>part.
> >Yes, I did understand that and said that from looking at the Solr
> >documentation, it looks like a lot of what we're currently doing is
> >(now?) possible with Lucene/Solr builtins as well.
> It seems not, your point was that *Solr* provides these things,

Yes, because at the beginning I was only comparing Solr to the existing
search server, but you are right, Solr is "only" a nice abstraction over
Lucene.

> and my point was that *Lucene* already provides this, and that Solr
> isn't adding anything to do this particular aspect it just provides
> access to what lucene offers.
> We may well be able to replace some of the custom filters with prebuilt ones
> that have come along later, but that is simply done by modifying the
> existing code.

With Solr, you'd just have to change a config file and add a test case,
which is imho an easier task for most people than digging into a large
codebase in any language.

> Glad you looked at these but I'm not going through them now however ...
> >
> >* MusicBrainzTokenizer
> >
> >   This one "splits words at punctuation characters, removing
> >   punctuation.  However, a dot that's not followed by whitespace is
> >   considered part of a token.  Splits words at hyphens, unless there's a
> >   number in the token, in which case the whole token is interpreted as a
> >   product number and is not split. Recognizes email addresses and
> >   internet hostnames as one token."
> >
> >   As I said in an earlier mail, this seems to be exactly what [5] is
> >   doing.
> >
> I did clearly explain to you why this was necessary and pointed you towards
> the bug tracker so unless I misunderstood you I don't understand why you
> still think there is no difference

Sorry, I'm honestly not sure why I didn't take your previous reply into
account when I wrote that mail :/

On the other hand, the ClassicTokenizer is what was formerly known as
the StandardTokenizer and your patch can be applied to it.

-- 
Wieland
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.