Re: Article about using snowball stemmer to dolanguage identificaction

iolalla <[email protected]>
Newsgroups gmane.comp.search.snowball
Message-ID <[email protected]>
Hi Martin,

Yes you are right what we do is to count removed endings,
but when we are talking about very similar languages, Italian,
Portuguese, Catalan, etc. we need another criteria and mixing
stopwords and Stemming gives very good results.

With Jan from cominvent we are analyzing the idea of using a
mix of ngrams (hunspell), Snowball and StopWords in order to decide.
In the next weeks we'll have a test environment to see which is the
best approach and will share with everybody the results and the
code.

Kind Regards.

El 30/06/2011 12:00, Martin Porter escribió:

Iolalla,

Very interesting. In the past I've used stopword lists for language
identification, and the results were adequate. But I did notice that
in Finnish there aren't so very many stopwords!

But I don't quite see how applying a stemmer leads to a language
identification. Do you count the number of valid endings removed, and
use that as a measure?

Like Cominvent, I wondered how the mixing was done in the hybrid approach,

Martin

---------------------------------------------------------------------------------

ISrael
Olalla chiCOte

[email protected]

#Linkedin http://es.linkedin.com/in/iolalla

#M +34 664 300 116

#T +34 948 102 408

Parque Tomás Caballero, 2, 6º 4ª

31006 Pamplona, España

iSOCO

enabling
the networked economy

www.isoco.com

P Please consider your environmental
responsibility before
printing this e-mail

_______________________________________________
Snowball-discuss mailing list
[email protected]
http://lists.tartarus.org/mailman/listinfo/snowball-discuss
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.