Re: [PATCH v3 3/4] KVM: TDX: Don't assume exit_reason[31:16] as all-0 in tdx_to_vmx_exit_reason()

Sean Christopherson <[email protected]>
Newsgroups dev.linux.lists.sashiko-reviews,org.kernel.vger.kvm
Message-ID <[email protected]>
On Tue, Aug 18, 2026, Xiaoyao Li wrote:
> On 8/18/2026 2:19 AM, Sean Christopherson wrote:
> > On Mon, Aug 17, 2026, Xiaoyao Li wrote:
> > > On 8/14/2026 11:12 PM, Sean Christopherson wrote:
> > > > > > And once we track the exit_reason separately from vp_enter_ret, IMO it all becomes
> > > > > > more logical and easier to follow.  vp_enter_ret holds the information about why
> > > > > > VP.ENTER returned/exited, while exit_reason holds information about why the _guest_
> > > > > > exited.  Obviously it's imperfect since we're still fudging EXIT_REASON_EPT_MISCONFIG,
> > > > > > but again, I think that's a better alternative than demuxing exit_reason in multiple
> > > > > > locations.
> > > > > 
> > > > > The question do we really need to swizzle EXIT_REASON_TDCALL to other exit
> > > > > reasons ahead? why cannot them just be handled in the central handler for
> > > > > EXIT_REASON_TDCALL?
> > > > 
> > > > Because as above, it's not as central as you think.  If we want to not swizzle
> > > > the exit_reason, then IMO the only sane way to do that is to not track exit_reason
> > > > for TDX vCPUs, i.e. move vcpu_vt.exit_reason back to vcpu_vmx and force TDX to
> > > > always demux vp_enter_ret every time.
> > > 
> > > How about adding a specific field to track the TDVMCALL leaf? Full diff as
> > > below (the EPT MISCONFIG part can be split into a separate one)
> > 
> > I like it even less than demuxing vp_enter_ret on demand.  It has all the same
> > flaws as demuxing vp_enter_ret, because it's effectively just a cache of the
> > result of demuxing vp_enter_ret.  And caching values on entry/exit boundaries
> > adds additional risk;  I can point you at a number of bugs in the past where KVM
> > consumed stale data, e.g. because some chunk of code got moved to run before a
> > cached field was refreshed.  That's unlikekly to be a problem here, but anytime
> > data is cached it introduces risk of consuming stale data.
> > 
> > And the cost of demuxing vp_enter_ret every time doesn't concern me, at all.
> > What I don't like is relying on call sites to know that the exit_reason needs to
> > be demuxed in the first place, because that will be brittle and error prone.
> 
> Do you mean that with my diff, the consumers of tdx_is_tdvmcall() rely on
> handle_tdcall() being called already?

That's one of my concerns, but it's not my biggest concern.  What I'm most worried
about is providing what is effectively an incomplete vcpu_vt.exit_reason, relative
to how vcpu_vt.exit_reason has been handled by KVM for years, and still is handled
by VMX.  It essentially violates the principle of least surprise: years of experience
and huge swaths of the code base will result in developers expecting exit_reason to
be _the_ source of truth.

There are obviously cases where additional state needs to be queried to handle the
exit, e.g. most visibly in handle_tdvmcall().  The big difference is that
handle_tdvmcall() only deals with "new", TDX-specific functionality, and it's
analgous to kvm_emulate_hypercall(), i.e. it behaves pretty much how the majority
of developers would expect it to behave.

> If so, I get your point.
> 
> > That's why I'd be ok if vcpu_vt.exit_reason simply didn't exist: it becomes
> > impossible to check the wrong exit field because there's only one such field.
> > 
> > I'm ok caching the fully processed vcpu_vt.exit_reason, i.e. with the code now,
> > because only the TDVMCALL path needs to be aware that it may have undergone
> > processing, *and* KVM can WARN if that processing didn't happen as expected.
> > 
> > Whereas checking the "right" exit reason requires doing so in multiple paths and
> > doesn't have a natural sanity check.
> 
> I'm OK with it.
> 
> After changing to use tdx_is_exit_reason_valid() instead of checking
> (*reason != -1u) in tdx_get_exit_info(), the only issue is trace_kvm_exit()
> prints 0x0000ffff when real EPT MISCONFIG happens. But it can be resolved by
> your suggestion below. So current implementation can still work.

Yep, that's one of the reasons why I don't want to overwrite the exit reason.
lmpx.com only provides a reader for public news (NNTP) servers. It is not affiliated with the servers or forums shown here and is not responsible for the content of articles, which is written by their respective authors.