Re: sed error message reports byte position instead of char position when program contains UTF-8
Camion SPAM <[email protected]>
| Newsgroups | gmane.comp.gnu.utils.bugs |
|---|---|
| Message-ID | <[email protected]> |
> Eli Zaretskii <[email protected]> wrote: > "Byte" is the only viable alternative, but that leaves the burden of > counting bytes on the user. If you use a very long sed script, this can be a problem. I happen to write sed scripts that are more than 1000 characters long. My last one is currently 1762 characters long and still growing. That's why I wrote a little bash function which would show me the n'th character. but this gave wrong positions on scripts with UTF-8 chars. The work-around is to change LC_CTYPE to C around the string processing part in my bash function, but, I believe that since sed supports multibytes characters, the error message should count characters and not bytes. btw : the error message states that the position is "char" and not "byte" : sed: -e expression #1, char 12: unknown option to `s'