Re: sed error message reports byte position instead of char position when program contains UTF-8

Camion SPAM <[email protected]>
Newsgroups gmane.comp.gnu.utils.bugs
Message-ID <[email protected]>
> Eli Zaretskii <[email protected]> wrote:

> "Byte" is the only viable alternative, but that leaves the burden of
> counting bytes on the user.

If you use a very long sed script, this can be a problem. 
I happen to write sed scripts that are more than 1000 characters long.
My last one is currently 1762 characters long and still growing.
That's why I wrote a little bash function which would show me 
the n'th character. but this gave wrong positions on scripts with
UTF-8 chars. 

The work-around is to change LC_CTYPE to C around the string 
processing part in my bash function, but, I believe that since sed 
supports multibytes characters, the error message should count 
characters and not bytes. btw : the error message states that the 
position is "char" and not "byte" : 

sed: -e expression #1, char 12: unknown option to `s'
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.