Weeksio's Progress this week
Jeff Weeks <[email protected]> Mon, 29 Jun 2015 15:00:48 -0400
| Newsgroups | gmane.comp.audio.musicbrainz.devel |
|---|---|
| Message-ID | <CA+V91tgyDdfEaOcgXs3FS3CUwcizvkBtZ6SGwQMQ76cYya6c_A@mail.gmail.com> |
--===============0565635425== Content-Type: multipart/alternative; boundary=001a11c30b4e3935ab0519acb5ba --001a11c30b4e3935ab0519acb5ba Content-Type: text/plain; charset=UTF-8 Progress this week: -Boosted Solr schema.xml version from 1.2 to 1.5. -Edited sir to reflect -I changed direction on how I'm porting the analysis from the old server to the new. Instead of using the old java ported as is to Solr, I'm determining the steps each analysis takes and defining those steps directly in schema.xml. This should make future maintenance much easier. -I've been making my way through each core deciding which of the new field types to use for each attribute, importing each core's data from the database, fixing errors as they pop up, trying to improve the indexing and think into the future a bit as to how we will utilize some of Solr's features. I'm though Annotation, Area, Artist, CDStub, Editor...and the further I go the quicker it goes as each one covers more cases likely to appear in future fields as I go. TODOs: -Porting the Area, Artist, Label Boost configurations -Add remaining entities to sir -Continue working through cores Questions: -MusicBrainzKeepAccentsAnalyzer: I'm still not sure why we need this...maybe I'm overlooking something. It looks like tokens aren't stored, only analyzed/indexed; but if we use the same analysis at query time as at index time (which it appears we do) the indexed tokens retaining their accents will never be accessed. ...correct? -_store fields: I really don't understand the purpose of these...could someone explain? -[SEARCH-371] <http://tickets.musicbrainz.org/browse/SEARCH-371> Would an ICUTransformFilter using Greek-Latin and Cyrillic-Latin rule sets do the trick? ...or are resources the primary concern? -I noticed our analyzers don't use stop words? Is this something to continue? Seems like a good conservative list of English stop words would be useful. Personal Criticism: -I've been really bad about getting my code up on github...hope to get that up today. I just haven't gotten into the habit of it....I'm so used to working by myself on school projects. Personal Observations: -Solr is awesome! It's super powerful...and it's really easy to get bogged down thinking about adding extras. Best, Jeff --001a11c30b4e3935ab0519acb5ba Content-Type: text/html; charset=UTF-8 Content-Transfer-Encoding: quoted-printable <div dir=3D"ltr"><div class=3D"gmail_default" style=3D"font-family:verdana,= sans-serif;font-size:small">Progress this week:</div><div class=3D"gmail_de= fault" style=3D"font-family:verdana,sans-serif;font-size:small"><br></div><= div class=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-si= ze:small">-Boosted Solr schema.xml version from 1.2 to 1.5.</div><div class= =3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:small"= >-Edited sir to reflect</div><div class=3D"gmail_default" style=3D"font-fam= ily:verdana,sans-serif;font-size:small">-I changed direction on how I'm= porting the analysis from the old server to the new. Instead of using the = old java ported as is to Solr, I'm determining the steps each analysis = takes and defining those steps directly in schema.xml.=C2=A0 This should ma= ke future maintenance much easier.=C2=A0</div><div class=3D"gmail_default" = style=3D"font-family:verdana,sans-serif;font-size:small">-I've been mak= ing my way through each core deciding which of the new field types to use f= or each attribute, importing each core's data from the database, fixing= errors as they pop up, trying to improve the indexing and think into the f= uture a bit as to how we will utilize some of Solr's features.=C2=A0 I&= #39;m though Annotation, Area, Artist, CDStub, Editor...and the further I g= o the quicker it goes as each one covers more cases likely to appear in fut= ure fields as I go.</div><div class=3D"gmail_default" style=3D"font-family:= verdana,sans-serif;font-size:small"><br></div><div class=3D"gmail_default" = style=3D"font-family:verdana,sans-serif;font-size:small">TODOs:</div><div c= lass=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:sm= all">-Porting the Area, Artist, Label Boost configurations</div><div class= =3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:small"= >-Add remaining entities to sir</div><div class=3D"gmail_default" style=3D"= font-family:verdana,sans-serif;font-size:small">-Continue working through c= ores</div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-se= rif;font-size:small"><br></div><div class=3D"gmail_default" style=3D"font-f= amily:verdana,sans-serif;font-size:small">Questions:</div><div class=3D"gma= il_default" style=3D"font-family:verdana,sans-serif;font-size:small"><br></= div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-serif;fo= nt-size:small">-MusicBrainzKeepAccentsAnalyzer:</div><div class=3D"gmail_de= fault" style=3D"font-family:verdana,sans-serif;font-size:small">I'm sti= ll not sure why we need this...maybe I'm overlooking something.=C2=A0 I= t looks like tokens aren't stored, only analyzed/indexed; but if we use= the same analysis at query time as at index time (which it appears we do) = the indexed tokens retaining their accents will never be accessed. =C2=A0..= .correct?</div><div class=3D"gmail_default" style=3D"font-family:verdana,sa= ns-serif;font-size:small"><br></div><div class=3D"gmail_default" style=3D"f= ont-family:verdana,sans-serif;font-size:small">-_store fields:</div><div cl= ass=3D"gmail_default" style=3D"font-family:verdana,sans-serif;font-size:sma= ll">I really don't understand the purpose of these...could someone expl= ain?</div><div class=3D"gmail_default" style=3D"font-family:verdana,sans-se= rif;font-size:small"><br></div><div class=3D"gmail_default" style>-<a href= =3D"http://tickets.musicbrainz.org/browse/SEARCH-371">[SEARCH-371]</a><br><= /div><div class=3D"gmail_default" style><p style=3D"margin:0px 0px 1em;padd= ing:0px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;fo= nt-size:12px;line-height:16.7999992370605px">Would an ICUTransformFilter us= ing Greek-Latin and Cyrillic-Latin rule sets do the trick?=C2=A0<span style= =3D"line-height:16.7999992370605px">...or are resources the primary concern= ?</span></p><p style=3D"margin:0px 0px 1em;padding:0px;color:rgb(0,0,0);fon= t-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;line-height:16.= 7999992370605px"><span style=3D"line-height:16.7999992370605px">-I noticed = our analyzers don't use stop words?=C2=A0 Is this something to continue= ?=C2=A0 Seems like a good conservative list of English stop words would be = useful.</span></p><p style=3D"margin:0px 0px 1em;padding:0px;color:rgb(0,0,= 0);font-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;line-heig= ht:16.7999992370605px">Personal Criticism:<br>-I've been really bad abo= ut getting my code up on github...hope to get that up today.=C2=A0 I just h= aven't gotten into the habit of it....I'm so used to working by mys= elf on school projects.</p><p style=3D"margin:0px 0px 1em;padding:0px;color= :rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;font-size:12px;= line-height:16.7999992370605px">Personal Observations:<br>-Solr is awesome!= =C2=A0 It's super powerful...and it's really easy to get bogged dow= n thinking about adding extras.</p><p style=3D"margin:0px 0px 1em;padding:0= px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-serif;font-si= ze:12px;line-height:16.7999992370605px"><br></p><p style=3D"margin:0px 0px = 1em;padding:0px;color:rgb(0,0,0);font-family:arial,FreeSans,Helvetica,sans-= serif;font-size:12px;line-height:16.7999992370605px"><span style=3D"line-he= ight:16.7999992370605px">Best,<br></span><span style=3D"line-height:16.7999= 992370605px">Jeff</span></p></div></div> --001a11c30b4e3935ab0519acb5ba-- --===============0565635425== Content-Type: text/plain; charset="us-ascii" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit Content-Disposition: inline _______________________________________________ MusicBrainz-devel mailing list [email protected] http://lists.musicbrainz.org/mailman/listinfo/musicbrainz-devel --===============0565635425==--