Weeksio's Progress this week

Jeff Weeks <[email protected]> Mon, 29 Jun 2015 15:00:48 -0400
Newsgroups gmane.comp.audio.musicbrainz.devel
Message-ID <CA+V91tgyDdfEaOcgXs3FS3CUwcizvkBtZ6SGwQMQ76cYya6c_A@mail.gmail.com>
--===============0565635425==
Content-Type: multipart/alternative; boundary=001a11c30b4e3935ab0519acb5ba

--001a11c30b4e3935ab0519acb5ba
Content-Type: text/plain; charset=UTF-8

Progress this week:

-Boosted Solr schema.xml version from 1.2 to 1.5.
-Edited sir to reflect
-I changed direction on how I'm porting the analysis from the old server to
the new. Instead of using the old java ported as is to Solr, I'm
determining the steps each analysis takes and defining those steps directly
in schema.xml.  This should make future maintenance much easier.
-I've been making my way through each core deciding which of the new field
types to use for each attribute, importing each core's data from the
database, fixing errors as they pop up, trying to improve the indexing and
think into the future a bit as to how we will utilize some of Solr's
features.  I'm though Annotation, Area, Artist, CDStub, Editor...and the
further I go the quicker it goes as each one covers more cases likely to
appear in future fields as I go.

TODOs:
-Porting the Area, Artist, Label Boost configurations
-Add remaining entities to sir
-Continue working through cores

Questions:

-MusicBrainzKeepAccentsAnalyzer:
I'm still not sure why we need this...maybe I'm overlooking something.  It
looks like tokens aren't stored, only analyzed/indexed; but if we use the
same analysis at query time as at index time (which it appears we do) the
indexed tokens retaining their accents will never be accessed.  ...correct?

-_store fields:
I really don't understand the purpose of these...could someone explain?

-[SEARCH-371] <http://tickets.musicbrainz.org/browse/SEARCH-371>

Would an ICUTransformFilter using Greek-Latin and Cyrillic-Latin rule sets
do the trick? ...or are resources the primary concern?

-I noticed our analyzers don't use stop words?  Is this something to
continue?  Seems like a good conservative list of English stop words would
be useful.

Personal Criticism:
-I've been really bad about getting my code up on github...hope to get that
up today.  I just haven't gotten into the habit of it....I'm so used to
working by myself on school projects.

Personal Observations:
-Solr is awesome!  It's super powerful...and it's really easy to get bogged
down thinking about adding extras.


Best,
Jeff

--001a11c30b4e3935ab0519acb5ba
Content-Type: text/html; charset=UTF-8
Content-Transfer-Encoding: quoted-printable

<div dir=3D"ltr"><div class=3D"gmail_default" style=3D"font-family:verdana,=
sans-serif;font-size:small">Progress this week:</div><div class=3D"gmail_de=
fault" style=3D"font-family:verdana,sans-serif;font-size:small"><br></div><=
div class=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-si=
ze:small">-Boosted Solr schema.xml version from 1.2 to 1.5.</div><div class=
=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:small"=
>-Edited sir to reflect</div><div class=3D"gmail_default" style=3D"font-fam=
ily:verdana,sans-serif;font-size:small">-I changed direction on how I&#39;m=
 porting the analysis from the old server to the new. Instead of using the =
old java ported as is to Solr, I&#39;m determining the steps each analysis =
takes and defining those steps directly in schema.xml.=C2=A0 This should ma=
ke future maintenance much easier.=C2=A0</div><div class=3D"gmail_default" =
style=3D"font-family:verdana,sans-serif;font-size:small">-I&#39;ve been mak=
ing my way through each core deciding which of the new field types to use f=
or each attribute, importing each core&#39;s data from the database, fixing=
 errors as they pop up, trying to improve the indexing and think into the f=
uture a bit as to how we will utilize some of Solr&#39;s features.=C2=A0 I&=
#39;m though Annotation, Area, Artist, CDStub, Editor...and the further I g=
o the quicker it goes as each one covers more cases likely to appear in fut=
ure fields as I go.</div><div class=3D"gmail_default" style=3D"font-family:=
verdana,sans-serif;font-size:small"><br></div><div class=3D"gmail_default" =
style=3D"font-family:verdana,sans-serif;font-size:small">TODOs:</div><div c=
lass=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:sm=
all">-Porting the Area, Artist, Label Boost configurations</div><div class=
=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:small"=
>-Add remaining entities to sir</div><div class=3D"gmail_default" style=3D"=
font-family:verdana,sans-serif;font-size:small">-Continue working through c=
ores</div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-se=
rif;font-size:small"><br></div><div class=3D"gmail_default" style=3D"font-f=
amily:verdana,sans-serif;font-size:small">Questions:</div><div class=3D"gma=
il_default" style=3D"font-family:verdana,sans-serif;font-size:small"><br></=
div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-serif;fo=
nt-size:small">-MusicBrainzKeepAccentsAnalyzer:</div><div class=3D"gmail_de=
fault" style=3D"font-family:verdana,sans-serif;font-size:small">I&#39;m sti=
ll not sure why we need this...maybe I&#39;m overlooking something.=C2=A0 I=
t looks like tokens aren&#39;t stored, only analyzed/indexed; but if we use=
 the same analysis at query time as at index time (which it appears we do) =
the indexed tokens retaining their accents will never be accessed. =C2=A0..=
.correct?</div><div class=3D"gmail_default" style=3D"font-family:verdana,sa=
ns-serif;font-size:small"><br></div><div class=3D"gmail_default" style=3D"f=
ont-family:verdana,sans-serif;font-size:small">-_store fields:</div><div cl=
ass=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:sma=
ll">I really don&#39;t understand the purpose of these...could someone expl=
ain?</div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-se=
rif;font-size:small"><br></div><div class=3D"gmail_default" style>-<a href=
=3D"http://tickets.musicbrainz.org/browse/SEARCH-371">[SEARCH-371]</a><br><=
/div><div class=3D"gmail_default" style><p style=3D"margin:0px 0px 1em;padd=
ing:0px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;fo=
nt-size:12px;line-height:16.7999992370605px">Would an ICUTransformFilter us=
ing Greek-Latin and Cyrillic-Latin rule sets do the trick?=C2=A0<span style=
=3D"line-height:16.7999992370605px">...or are resources the primary concern=
?</span></p><p style=3D"margin:0px 0px 1em;padding:0px;color:rgb(0,0,0);fon=
t-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;line-height:16.=
7999992370605px"><span style=3D"line-height:16.7999992370605px">-I noticed =
our analyzers don&#39;t use stop words?=C2=A0 Is this something to continue=
?=C2=A0 Seems like a good conservative list of English stop words would be =
useful.</span></p><p style=3D"margin:0px 0px 1em;padding:0px;color:rgb(0,0,=
0);font-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;line-heig=
ht:16.7999992370605px">Personal Criticism:<br>-I&#39;ve been really bad abo=
ut getting my code up on github...hope to get that up today.=C2=A0 I just h=
aven&#39;t gotten into the habit of it....I&#39;m so used to working by mys=
elf on school projects.</p><p style=3D"margin:0px 0px 1em;padding:0px;color=
:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;=
line-height:16.7999992370605px">Personal Observations:<br>-Solr is awesome!=
=C2=A0 It&#39;s super powerful...and it&#39;s really easy to get bogged dow=
n thinking about adding extras.</p><p style=3D"margin:0px 0px 1em;padding:0=
px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;font-si=
ze:12px;line-height:16.7999992370605px"><br></p><p style=3D"margin:0px 0px =
1em;padding:0px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-=
serif;font-size:12px;line-height:16.7999992370605px"><span style=3D"line-he=
ight:16.7999992370605px">Best,<br></span><span style=3D"line-height:16.7999=
992370605px">Jeff</span></p></div></div>

--001a11c30b4e3935ab0519acb5ba--


--===============0565635425==
Content-Type: text/plain; charset="us-ascii"
MIME-Version: 1.0
Content-Transfer-Encoding: 7bit
Content-Disposition: inline

_______________________________________________
MusicBrainz-devel mailing list
[email protected]
http://lists.musicbrainz.org/mailman/listinfo/musicbrainz-devel
--===============0565635425==--