Re: Article about using snowball stemmer to dolanguage identificaction
iolalla <[email protected]>
| Newsgroups | gmane.comp.search.snowball |
|---|---|
| Message-ID | <[email protected]> |
Hi Martin, Yes you are right what we do is to count removed endings, but when we are talking about very similar languages, Italian, Portuguese, Catalan, etc. we need another criteria and mixing stopwords and Stemming gives very good results. With Jan from cominvent we are analyzing the idea of using a mix of ngrams (hunspell), Snowball and StopWords in order to decide. In the next weeks we'll have a test environment to see which is the best approach and will share with everybody the results and the code. Kind Regards. El 30/06/2011 12:00, Martin Porter escribió: Iolalla, Very interesting. In the past I've used stopword lists for language identification, and the results were adequate. But I did notice that in Finnish there aren't so very many stopwords! But I don't quite see how applying a stemmer leads to a language identification. Do you count the number of valid endings removed, and use that as a measure? Like Cominvent, I wondered how the mixing was done in the hybrid approach, Martin --------------------------------------------------------------------------------- ISrael Olalla chiCOte [email protected] #Linkedin http://es.linkedin.com/in/iolalla #M +34 664 300 116 #T +34 948 102 408 Parque Tomás Caballero, 2, 6º 4ª 31006 Pamplona, España iSOCO enabling the networked economy www.isoco.com P Please consider your environmental responsibility before printing this e-mail _______________________________________________ Snowball-discuss mailing list [email protected] http://lists.tartarus.org/mailman/listinfo/snowball-discuss