Edit report at https://bugs.php.net/bug.php?id=74312&edit=1
ID: 74312
Updated by: [email protected]
Reported by: furun at arcor dot de
Summary: Filter BOMs out of PHP output
Status: Re-Opened
Type: Bug
Package: Unicode Engine related
Operating System: Win7
PHP Version: 7.0.17
Block user comment: N
Private report: N
New Comment:
Lets not start calling people idiots...
Here is a semi-recent internals thread on the topic of BOM handling: http://markmail.org/message/besjw22hxlpwlvdh
My opinion here is essentially this:
a) First, a philosophical point. PHP, in a default configuration, does not care about file encoding. As long as it's ASCII compatible, PHP does not care whether your file uses UTF-8, ISO-8859-1 or Windows-1251. PHP does not perform any validation or conversion, it gives you everything back exactly as it received it. There *is* a special mode in which PHP does all these things, and that's zend.multibyte. If this mode is enabled, PHP can convert character encodings, strip BOMs, etc. Of course, we know that in practice there is very little interest in this.
b) Second, a practical point. You are considering the case where the BOM was accidentally introduced by bad tooling (say, people writing PHP code in MS Word). However, it may also be introduced *intentionally*, to be produced as part of the output. (For the sake of an example, assume you have template files including BOMs, because you want to deliver a response with BOMs, as tooling on the consuming side requires it.) If nothing else, this is a backwards compatibility break and a rather subtle one at that. As such, such a change could probably only be made behind an ini setting, which would probably not help with your particular concern.
Previous Comments:
------------------------------------------------------------------------
[2017-04-01 15:02:21] furun at arcor dot de
(spam2) I understand you there, but what is the easier way to do, to reeducate the entire planet, or to take the stone out of the way.
People use there favorite tools, and unexperienced developers are publishing there first plugins/codes for other software, user using it without possibility to proof, and professionals make mistakes or search bugs on the wrong end (they then just not need forum help).
My arguments are there to held the php community which has to deal with it, and passive users with no coding background, and time spend and wait for bug fixes, wen it goes wrong, an it will go wrong.
You as a experiences developer will not feel any harm anyhow, if BOM and ?> is cleaned, or not.
My effort here is for the "wild real-life, all out there".
------------------------------------------------------------------------
[2017-04-01 14:08:06] spam2 at rhsoft dot net
and problems have to be fixed where they are
* idiotic editors which add BOM to non-UTF16 files
* idiotic php editors hich add newlines at the end of the file
period
------------------------------------------------------------------------
[2017-04-01 14:05:59] spam2 at rhsoft dot net
> This is rhetoric to protect a dogma, not a technical argument
franly, if you did not realize it - i am not a php upstream-developer and so don't need to protect any dogma - i am just a userland php-script developer for 15 years now, have written some hundret thousand lines of php-code and frankly tell you if that are real issues for a developer he RELLAY should stop develop software at all
------------------------------------------------------------------------
[2017-04-01 13:40:27] furun at arcor dot de
This is rhetoric to protect a dogma, not a technical argument. If a syntax definition is acting like a pitfall, it should be changed.
Bug prevention is the contrary of randomness. Human errors are random, and bug prevention takes them out of the way.
The write development strategies is to see, how many time is a ending newline/empty-char intended, and how many time coders are stumble over it unintended. And then take unwanted randomness out of the no-game, if indented newlines/empty-chars are very rare and can be done more explicit.
In the middle of a PHP and HTML mixed script, precision for <?php ?> delimitations makes sense, and should not be faded.
Tanks that there is no enforcement to use ending ?>, they are useless anyway, specially in near 100% php code scripts.
Smart coders don't write any ending ?> to safe there time.
But i see the empty-char after last ?> as not so buggy like the BOM, but i miss the good infos and experiences to decide what is best.
A very active php community forum helper maybe can tell, how often this becomes a pitfall, and if it makes sense to chance definitions.
(Just frequenting forums sometime, i seen it many time, ?>.. seems to be a pitfall for me.)
------------------------------------------------------------------------
[2017-04-01 02:40:31] spam2 at rhsoft dot net
> And yes, why not do a implicit trim()
because when i want a gambling machine i use a gambling machine and not a programming language
------------------------------------------------------------------------
The remainder of the comments for this report are too long. To view
the rest of the comments, please view the bug report online at
https://bugs.php.net/bug.php?id=74312
--
Edit this bug report at https://bugs.php.net/bug.php?id=74312&edit=1
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.