Re: Chinese and Japanese stemmeing algorithm

Richard Boulton <[email protected]>
Newsgroups gmane.comp.search.snowball
Message-ID <CAGb4Wn__aoedG-sxXSzS0R56=tucDbHfknygTfm56vU=dgj4PQ@mail.gmail.com>
On 2 November 2011 10:51, Martin Porter <[email protected]> wrote:
> Does anyone else contributing to snowball
> discuss have more knowledge on what is currently available?

There are a few things available.

One easy approach is to use bi-grams of CJK characters; which doesn't
work wornderfully, but is better than nothing.  There are some bits of
code lying around to assist with that; for example,
http://code.google.com/p/cjk-tokenizer/

There are also more sophisticated approaches, generally involving some
use of dictionaries.  I don't know of standalone code for doing these,
but we had a Google-Summer-of-Code student with Xapian this year who
implemented quite a lot of stuff for Chinese word segmentation; his
work hasn't been integrated into Xapian core yet, but the trac page
describing it (with links to the code) is
http://trac.xapian.org/wiki/GSoC2011/ChineseSegmentationAnalysis

Olly Betts was his primary mentor, so may be able to give more detail.

-- 
Richard
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.