Fw: Repurposing U+2019 RIGHT SINGLE QUOTATION MARK as a Lexical Word Divider for the SE Asian scripts that have NO SPACE BETWEEN WORDS

David Haslam <dfhdfh-g/[email protected]> Thu, 29 May 2025 14:46:57 +0000
Newsgroups gmane.comp.literature.sword.devel
Message-ID <OpNjpLXunNiVcE50IRqXpBrhxLMuE49JY9E-gHIa9rF80paaNNG9LXWs5rg_Aj8yMNW4eL5GVuKW9x7HENC-yFp77h3CvDjXsAUixPpM6qs=@protonmail.com>
NB. I have cancelled the earlier email because the attachment was too large for sword-devel.
It had been in the queue for moderator approval.

The eXperimental module KhmerNTx.zip may now be downloaded from this [link](https://app.box.com/s/e613wf1qdxbjmvux9gbb6vmes33d2rol) on my box.net account.

Please see below for the significant details.

Best regards,

David

Sent with [Proton Mail](https://pr.tn/ref/SWXT9A5YZ67G) secure email.

------- Forwarded Message -------
From: David Haslam <[email protected]>
Date: On Thursday, May 29th, 2025 at 9:26 AM
Subject: Repurposing U+2019 RIGHT SINGLE QUOTATION MARK as a Lexical Word Divider for the SE Asian scripts that have NO SPACE BETWEEN WORDS
To: sword-devel mailing list <[email protected]>
CC: [email protected] <[email protected]>, Modules Issues <[email protected]>

> Dear SWORD Developers (and our Modules Team),
>
> While watching the [livestream funeral](https://www.youtube.com/live/zC4hXOgqBak?si=JZ7JiM7j_fHW-sQl) of OT Scholar the late Gordon D Wenham yesterday (St Mary's Church, Charlton Kings), I had a bright idea.
>
> I'd been working recently on potential improvements for the KhmerNT module relating to marking the Lexical Word Divisions.
> Khmer is one of the languages of SE Asia whose Writing System (aka Script) largely has NO SPACE BETWEEN WORDS.
> Others include: Lao, Thai, Myanmar (aka Burmese), together with other languages in the region that employ one of these scripts (e.g. Isaan).
>
> Until the present, the KhmerNT module makes use of the ZWSP = Zero Width Space to mark lexical word boundaries.
> This helps with SWORD search for whole words, because even though the divisions between words are invisible to human eyes, they are accessible to computer software.
>
> Wouldn't it be nice if ... (cue to sing the melody by the Beach Boys) 🎢
>
> - We could instead use a visible Unicode character
> - That character could be hidden by means of an existing SWORD filter
>
> There is such a character!!!
>
> - U+2019 is one of the codepoints hidden (or changed) by the filter UTF8GreekAccents.
>
>> U+2019 (RIGHT SINGLE QUOTATION MARK) is commonly used in digital editions of the NT Greek as the apostrophe, not as a quotation mark.
>>
>> In NT Greek, it appears in:
>>
>> - Elisions: When a vowel at the end of a word is dropped (e.g., δι’ instead of διά before a vowel).
>> - Contractions or abbreviations: e.g., ἐπ’ for ἐπί, καθ’ for κατά.
>> While U+2019 is typographically correct for apostrophes in modern typesetting, some older or simpler digital texts may use U+0027 (straight apostrophe). However, U+2019 is the preferred character in high-quality, properly typeset Greek texts.
>
> I then set about to test my idea by making a further update to an already eXperimental version of the module, provisionally named KhmerNTx.
>
> It "worked like a dream". 😎
>
> With Greek accents hidden, the text looks like this:
>
>> αžαŸ’αž‰αž»αŸ†αž–αŸαžαŸ’αžšαž»αžŸ αž‡αžΆαžŸαžΆαžœαž€αžšαž”αžŸαŸ‹αž–αŸ’αžšαŸ‡αž™αŸαžŸαŸŠαžΌαž‚αŸ’αžšαž·αžŸαŸ’αžŠ αž‡αžΌαž“αž…αŸ†αž–αŸ„αŸ‡αž–αž½αž€αž’αŸ’αž“αž€αžŠαŸ‚αž›αž–αŸ’αžšαŸ‡αž‡αžΆαž˜αŸ’αž…αžΆαžŸαŸ‹αž”αžΆαž“αž‡αŸ’αžšαžΎαžŸαžšαžΎαžŸ αž αžΎαž™αžŠαŸ‚αž›αž”αžΆαž“αž”αŸ‚αž€αžαŸ’αž‰αŸ‚αž€αž‚αŸ’αž“αžΆαž‘αŸ…αžŸαŸ’αž“αžΆαž€αŸ‹αž“αŸ…αž”αžŽαŸ’αžŠαŸ„αŸ‡αž’αžΆαžŸαž“αŸ’αž“αž“αŸ…αžŸαŸ’αžšαž»αž€αž”αŸ‰αž»αž“αžαž»αžŸ αžŸαŸ’αžšαž»αž€αž€αžΆαž‘αžΆαž‘αžΈ αžŸαŸ’αžšαž»αž€αž€αžΆαž”αŸ‰αžΆαžŠαžΌαž‚αžΆ αžŸαŸ’αžšαž»αž€αž’αžΆαžŸαŸŠαžΈ αž“αž·αž„αžŸαŸ’αžšαž»αž€αž”αŸ‰αžΈαž’αžΌαž“αžΆ (I Peter 1:1 [KhmerNTx])
>
> With Greek accents displayed, the text looks like this:
>
>> αžαŸ’αž‰αž»αŸ†β€™αž–αŸαžαŸ’αžšαž»αžŸ αž‡αžΆβ€™αžŸαžΆαžœαž€β€™αžšαž”αžŸαŸ‹β€™αž–αŸ’αžšαŸ‡β€™αž™αŸαžŸαŸŠαžΌβ€™αž‚αŸ’αžšαž·αžŸαŸ’αžŠ αž‡αžΌαž“β€™αž…αŸ†αž–αŸ„αŸ‡β€™αž–αž½αž€αž’αŸ’αž“αž€β€™αžŠαŸ‚αž›β€™αž–αŸ’αžšαŸ‡αž‡αžΆαž˜αŸ’αž…αžΆαžŸαŸ‹β€™αž”αžΆαž“β€™αž‡αŸ’αžšαžΎαžŸαžšαžΎαžŸ αž αžΎαž™β€™αžŠαŸ‚αž›β€™αž”αžΆαž“β€™αž”αŸ‚αž€αžαŸ’αž‰αŸ‚αž€β€™αž‚αŸ’αž“αžΆβ€™αž‘αŸ…β€™αžŸαŸ’αž“αžΆαž€αŸ‹β€™αž“αŸ…β€™αž”αžŽαŸ’αžŠαŸ„αŸ‡αž’αžΆαžŸαž“αŸ’αž“β€™αž“αŸ…β€™αžŸαŸ’αžšαž»αž€β€™αž”αŸ‰αž»αž“αžαž»αžŸ αžŸαŸ’αžšαž»αž€β€™αž€αžΆαž‘αžΆαž‘αžΈ αžŸαŸ’αžšαž»αž€β€™αž€αžΆαž”αŸ‰αžΆαžŠαžΌαž‚αžΆ αžŸαŸ’αžšαž»αž€β€™αž’αžΆαžŸαŸŠαžΈ αž“αž·αž„β€™αžŸαŸ’αžšαž»αž€β€™αž”αŸ‰αžΈαž’αžΌαž“αžΆ (I Peter 1:1 [KhmerNTx])
>
> I have attached the compressed module for any of you to explore & play with further.
>
> Aside: The previous update already made use of the OSIS XML w element to enclose each lexical Khmer word. That remains the case.
> In this way, the module source text is ready to be adaptedforfurther enhancements such as adding Strong's numbers, etc, to make a Study Edition.
>
> Steve Hyde and the translators in Cambodia are currently preparing to publish the complete Khmer Bible.
> He has requested my assistance in improving the actual word divisions for the 39 OT books.
> I've already been sent the source text, exported from their database.
>
> Since early May, I have been exploring how the Grok AI engine can make a positive contribution to the success of this challenging task.
> More on that subject later.
>
> Best regards,
>
> David
>
> Sent with [Proton Mail](https://pr.tn/ref/SWXT9A5YZ67G) secure email.

_______________________________________________
sword-devel mailing list: [email protected]
http://crosswire.org/mailman/listinfo/sword-devel
Instructions to unsubscribe/change your settings at above page