[php-src] Issue #23050: Segfault in `execute_ex`: `EX(run_time_cache)` is NULL for a trait-copied `private static` method (no opcache, no JIT)
[email protected] (peanutgraphic)
| Newsgroups | php.bugs |
|---|---|
| Message-ID | <N2VP6p3bCV6wQ8kFeHHSrcrKUwYxtoXdfIiRw6PXV7I@main.internal.php.net> |
Issue: https://github.com/php/php-src/issues/23050
Author: peanutgraphic
### Description
We are seeing a reproducible-in-aggregate segfault in the Zend VM on PHP 8.3.31
and 8.3.32 (Ubuntu 24.04, deb.sury.org builds, NTS). It occurs ~36 times/week on
one CI host under a Laravel application served by `php -S` (via
`php artisan serve`). Because the built-in server is a single process, each crash
takes the server down and every subsequent request is refused.
**The engine executes an op_array whose `run_time_cache` is NULL, then performs a
run-time cache slot lookup against it.**
Faulting instruction and registers, from a core dump (8.3.32):
```
=> execute_ex+15353: mov (%rdx,%rax,1),%rbx
rdx = 0x0 <- run-time cache base
rax = 0x8 <- cache slot offset
execute_ex+15357: test %rbx,%rbx
execute_ex+15360: je ...
```
Read directly from the core:
```
run_time_cache (execute_data + 0x40) = 0x0
```
Kernel log line (identical every time, same binary offset on every occurrence):
```
php8.3[3641252]: segfault at 8 ip ...2a3729 sp ... error 4 in php8.3[...+32f000]
```
`error 4` = user-mode read of a non-present page; the faulting address is
literally `0 + 8`, i.e. slot 1 of a NULL cache.
### The function involved
Frame 0 is always the same PHP function, reached from different callers. Walking
`execute_data` out of the core (`%r14`; `func` at `+0x18`, `prev_execute_data` at
`+0x30`, `function_name` at `+0x08`, `zend_string.val` at `+0x18`):
```
#0 resolveTenantId <- always frame 0
#1 {closure} (an Eloquent global scope closure)
#2 {closure} -> callScope -> applyScopes
core A: ... -> get -> first -> resolveRouteBinding -> substituteImplicitBindings
core B: ... -> toBase -> __call -> currentAgencyHasStudio -> {closure}
```
`resolveTenantId()` is a **`private static` method defined in a trait**, and that
trait is used by ~68 classes — so the method exists as ~68 separate op_arrays,
each with its own `ZEND_MAP_PTR` run-time cache. It is called on a very hot path
(every scoped query). Its body is small:
```php
private static function resolveTenantId(): int|string|null
{
if (! function_exists('app')) {
return null;
}
$resolver = app()->bound(TenantResolver::class)
? app(TenantResolver::class)
: null;
return $resolver?->currentTenantId();
}
```
### Ruled out
- **opcache / JIT** — `opcache.enable_cli => Off`. opcache.so is loaded as a
module but caches nothing in this SAPI; JIT therefore never engages. (This
distinguishes it from #22158, which is JIT-specific, and #9396, which requires
opcache.)
- **`PHP_CLI_SERVER_WORKERS`** — unset; confirmed absent from the core's
environment, so the experimental forking server is not involved.
- **pcov / xdebug** or any other VM-hooking extension — not among the 65
extensions loaded in the crashed process.
- **Userland recursion into the same scope** — the class queried by the resolver
is exempt from the global scope, so there is no re-entry loop.
- **PHP 8.3.33** — its NEWS contains only Date/GD/PGSQL/Phar entries; nothing
engine-related, so we do not expect it to change this.
### Environment
- PHP 8.3.32 (`8.3.32-1+ubuntu24.04.1+deb.sury.org+1`) and 8.3.31, both NTS, both
affected, on two separate hosts
- Ubuntu 24.04
- SAPI: `cli-server` (`php -S`)
- opcache loaded, `opcache.enable_cli=Off`, JIT inactive
- 65 extensions loaded
### What I have and what I don't
I have core dumps and can extract any additional structure from them on request
(op_array fields, the `zend_function` at `execute_data+0x18`, the opline, etc.).
**I do not have a minimal reproducer, and I have tried.** The crash is frequent
in aggregate (~5/day) but not deterministic per-request, which is consistent with
an allocation- or lifetime-related issue around the run-time cache rather than a
specific input.
Attempts, both run on the *same* affected binary and host, both **negative**:
1. Structural mimic — one trait with a `private static` method calling
`function_exists()`, a container-ish helper and a nullsafe method call; 68
classes using the trait; each invoked through a `static function` closure via
`call_user_func` (to mirror the global-scope indirection); served by
`php -S`. **400 requests, ~186k calls: no crash.**
2. Same, scaled to **1500 trait-using classes + 1500 free functions** to force
`ZEND_MAP_PTR` table growth and reallocation across a long-lived
`php -S` process. **500 requests: no crash.**
So the trait-copy structure *alone* is not sufficient. Something else in the real
workload is required — plausibly overall arena/memory pressure, the volume of
classes autoloaded per request, or live PDO/session activity. I mention the
negative results explicitly so nobody re-runs the same experiment.
I have the core dumps and can extract any additional structure from them on
request, or run a diagnostic/assertion build against the live workload — that is
likely the fastest route to a fix, since the workload reproduces it reliably in
aggregate even though a reduced case does not.