Re: [PATCH v16 07/45] arm64: mm: Handle Granule Protection Faults (GPFs)
Catalin Marinas <[email protected]>
| Newsgroups | dev.linux.lists.linux-coco,dev.linux.lists.kvmarm,org.infradead.lists.linux-arm-kernel,org.kernel.vger.kvm,org.kernel.vger.linux-kernel |
|---|---|
| Message-ID | <[email protected]> |
On Wed, Aug 12, 2026 at 06:12:09PM +0530, Pavan Kondeti wrote: > On Tue, Aug 11, 2026 at 04:11:13PM +0100, Suzuki K Poulose wrote: > > On 11/08/2026 15:44, Catalin Marinas wrote: > > > On Mon, Aug 03, 2026 at 02:43:23PM +0100, Steven Price wrote: > > > > If the host attempts to access granules that have been delegated for use > > > > in a realm these accesses will be caught and will trigger a Granule > > > > Protection Fault (GPF). > > > > > > > > A fault during a page walk signals a bug in the kernel and is handled by > > > > oopsing the kernel. A non-page walk fault could be caused by user space > > > > having access to a page which has been delegated to the kernel and will > > > > trigger a SIGBUS to allow debugging why user space is trying to access a > > > > delegated page. > > > > > > > > Reviewed-by: Suzuki K Poulose <[email protected]> > > > > Reviewed-by: Gavin Shan <[email protected]> > > > > Signed-off-by: Steven Price <[email protected]> > > > > --- > > > > Changes since v10: > > > > * Don't call arm64_notify_die() in do_gpf() but simply return 1. > > > > Changes since v2: > > > > * Include missing "Granule Protection Fault at level -1" > > > > --- > > > > arch/arm64/mm/fault.c | 28 ++++++++++++++++++++++------ > > > > 1 file changed, 22 insertions(+), 6 deletions(-) > > > > > > > > diff --git a/arch/arm64/mm/fault.c b/arch/arm64/mm/fault.c > > > > index 85e23388f9bb..ea3ae0ca7dba 100644 > > > > --- a/arch/arm64/mm/fault.c > > > > +++ b/arch/arm64/mm/fault.c > > > > @@ -909,6 +909,22 @@ static int do_tag_check_fault(unsigned long far, unsigned long esr, > > > > return 0; > > > > } > > > > +static int do_gpf_ptw(unsigned long far, unsigned long esr, struct pt_regs *regs) > > > > +{ > > > > + const struct fault_info *inf = esr_to_fault_info(esr); > > > > + > > > > + die_kernel_fault(inf->name, far, esr, regs); > > > > + return 0; > > > > +} > > > > + > > > > +static int do_gpf(unsigned long far, unsigned long esr, struct pt_regs *regs) > > > > +{ > > > > + if (!is_el1_instruction_abort(esr) && fixup_exception(regs, esr)) > > > > + return 0; > > > > + > > > > + return 1; > > > > +} > > > > > > Is there a valid case for fixup_exception() here? IOW, do we ever have a > > > valid user mapping of the pages delegated to a guest? If the above is > > > considered a kernel bug, I'd not silently ignore this (like return less > > > bytes copied or -EFAULT to user) but rather warn, potentially > > > rate-limited. > > > > Good question. This shouldn't be a valid case. We expect the VMM to use > > guest_memfd and that should prevent any mmaps and thus fixups shouldn't > > be required. That said, this series doesn't enforce that the VMM uses > > GMEM backed memslots for Guest RAM. We should probably do that while > > mapping things in. > > > > With that, we could drop that fixup and scream a bit. > > There is a valid case for fixup_exception() to be needed here in GPF > handling. > > -000 |load_unaligned_zeropad(inline) > -000 |hash_name(inline) > -000 |link_path_walk() > -001 |path_lookupat() > -002 |filename_lookup() > -003 |vfs_statx() > -004 |vfs_fstatat() > > We observed this in Android running Gunyah when the page is mapped in > EL1 but unmapped at EL2. path_lookupat() can actually handle this > via fixup_exeption() when a word load crosses the page boundary. > However, Gunyah injects a Synchronous External Abort and we have > a downstream patch [1] that adds fixup_exception() in do_sea(). pKVM > injects [2] such faults back to EL1 and fixup_exception() is taken care. Ah, good point, completely forgot about load_unaligned_zeropad(). Since we don't unmap the linear map for delegated pages, we'll need the fixup_exception(). And I guess warning in this case is not desirable either. We could limit it to EX_TYPE_KACCESS_ERR_ZERO and EX_TYPE_LOAD_UNALIGNED_ZEROPAD, though not sure it's worth it. -- Catalin