Internet-Draft draft-das-rats-frontier-model-extraction-00.txt is now
available.
Title: Beyond Attestation: Hardware-Enforced Protection of OpenAI and Anthropic Claude Models Against Unauthorized Extraction and Distillation
Author: Sangam Das
Name: draft-das-rats-frontier-model-extraction-00.txt
Pages: 26
Dates: 2026-08-31
Abstract:
As frontier AI providers such as Anthropic and OpenAI expose
increasingly rich inference interfaces to partners, evaluators, and
enterprise tenants, richer-than-normal outputs -- for example logits,
log-probability distributions, embeddings, hidden states,
intermediate activations, KV-cache material, and diagnostic state --
become practical extraction vectors rather than theoretical concerns.
In this document, a "release" is any point at which such information
leaves a provider's protected execution environment and becomes
usable outside it, whether returned directly to a caller, cached,
forwarded, or otherwise materialized. Existing authentication,
confidential-computing, and remote-attestation mechanisms can
establish important properties of the requester and execution
environment, but they do not by themselves establish that each
specific release of sensitive model information is currently
authorized to cross that protected-to-unprotected boundary.
This document describes that gap as a pre-effectuation release-
control problem. It presents an execution-finality architecture in
which computation is separated from authority to release. A
sensitive result is treated as a Candidate Release, remains non-
effective outside the protected domain, and becomes externally
available only after release-specific validation, rollback-resistant
extraction-state evaluation, bounded authorization, and verification
at a controlled Finality Sink. The document discusses anti-bypass
requirements, remote-attestation composition, latency, legacy
deployment, industrial relevance, and a concrete model-extraction
scenario.
The mechanism is not claimed to prevent all forms of model
distillation. Its narrower objective is to reduce unauthorized or
excessive extraction of privileged model information that can
materially improve model stealing, reconstruction, imitation, or
distillation. Because release authority is bound to rollback-
resistant, atomically consumed extraction state rather than to a
bearer credential, the volume of privileged teacher signal an
adversary can accumulate through the governed path is capped by
authorized state transitions, not by request volume or attacker
persistence alone.
The IETF datatracker status page for this Internet-Draft is:
https://datatracker.ietf.org/doc/draft-das-rats-frontier-model-extraction/
There is also an HTML version available at:
https://www.ietf.org/archive/id/draft-das-rats-frontier-model-extraction-00.html
Internet-Drafts are also available by rsync at:
rsync.ietf.org::internet-drafts
_______________________________________________
I-D-Announce mailing list -- [email protected]
To unsubscribe send an email to [email protected]
lmpx.com only provides a reader for public news (NNTP) servers. It is not
affiliated with the servers or forums shown here and is not responsible for
the content of articles, which is written by their respective authors.