Re: Per-page statistics

Ray Saintonge <[email protected]> Thu, 25 Dec 2003 14:59:36 -0800
Newsgroups gmane.science.linguistics.wikipedia.international
Message-ID <[email protected]>
Ramanan Selvaratnam wrote:

>On Wed, 2003-12-24 at 14:43, Roozbeh Pournader wrote:
>
>>b) Sometimes a single German word contains information equivalent to
>>three English words. Putting them next to each other will suggest people
>>to compare them, which is just idiotic.
>>
>In the case of Tamil this is a very good analogy except joining words
>are governed by some grammatical rules which seems to be getting
>extinct.
>
>The concept of full stop itself was introduced only around mid 19th
>century by the person who helped to translate the Bible for a
>missionary. 
>
>So sometimes the word count for the same article after an edit in the
>wiki might differ but not the substance.
>
One should deal with the question in the spirit in which it was asked. 
 It was trying to find a measure for the relative completeness of the 
versions of an article in two different languages.  For reasons that 
have been raised using words as the basis for such a measure would 
simply not work.  

We can, however, measure the relative sizes of two text files.  We can 
also develop a statistical measure to compare the sizes of two texts. 
 The Rosetta Project uses the first 3 chapters of the Bible's book of 
Genesis as a standard, and provides version in more than 1,000 
languages.  A bare text in any of these languages will have a certain 
length.  Comparing that with the text in another language can give us a 
factor that we can then apply to one of our text to give a statistical 
estimate of how long a translated text should be in another language.

Ec