Fetching from the wire…
Public story · 2026-09-12 · high
A study of 115 trusted-execution setups found most can't prove their attested code matches public source, and only one developer called that a priority.
Why now: The paper posting to arXiv in September 2026 is the timing hook: it's dated 2609.11411, the paper's own arXiv identifier.
Remote attestation is supposed to let you verify that hardware is running the exact code it claims to run, not code you're trusting on faith. A study of 115 real-world deployments across Intel SGX, Intel TDX, and AMD SEV found that 91% of them can't back that up, and 80% didn't publish both the source code and a reference build a verifier could check against, per the paper.
Without a reproducible build, an attestation measurement is just a hash you're told to believe, with no way to independently confirm it matches any specific, inspectable source. That's the entire mechanism attestation is supposed to provide, and for the vast majority of deployments studied, it's missing.
The researchers didn't stop at scanning artifacts. They contacted maintainers of 50 SGX projects and interviewed 12 developers directly. Exactly one said reproducibility was a development priority, a number that turns a measurement gap into a statement about intent: most teams aren't attempting reproducibility at all.
If reproducibility barely registers as a priority for the people building these systems, buyers evaluating a TEE-based product on its attestation claims are mostly evaluating a claim nobody checked. Anyone relying on attestation as a security boundary, in a supply chain or a confidential-computing product, should ask the vendor directly whether they publish a reference build and matching source, because the paper's numbers say most won't have one ready.
Watch whether any TEE vendor responds by publishing reproducible build pipelines as a selling point. If that 9% figure doesn't move over the next year, it confirms the gap is structural, not an oversight waiting to be fixed.
Each link below shares sources, entities, or timing with this story.
Researchers loaded five systems with a revoked policy and its replacement, then measured retrieval and downstream action across nine policy scenarios, nine models and six defense conditions. Wherever the revocation label was visible to the retrieval layer, the revoked fact cam...
The study extracted 130 clean atomic state transitions from 707 real issues in SWE-bench Lite and Verified. Plain RAG scored 0.57-0.59 answer accuracy; an LLM reranker didn't help and added latency, about 18 seconds against 2.1. A (subject, relation, object) supersession memor...
Three LLM review systems compared within-subjects: detailed explanation plus feedback, feedback only, no explanation. Full explanations produced the highest perceived trust (M = 3.99/5) but not the highest agreement; moderate explanation won agreement at 89.22%. More explanati...
A July 2 evaluation tested semantic chunking against simple approaches on long structured academic theses using RAGAs, and the sophisticated method didn't win. Performance varied more with document formatting, preprocessing, and query type than with chunking strategy. The auth...
The authors extract a steering direction from the model's existing tool-use preference signal and apply it at inference, producing monotonic control over how often the agent reaches for a tool while keeping invocations valid (arXiv 2608.25198). Open-domain QA accuracy with liv...
Repeat-After-Me is a black-box adaptive visual prompt injection reaching above 80% attack success on Qwen3.6-27B and 47% on GPT-5.5, under a realistic setting where the benign user prompt is unrelated to and doesn't authorize the injected task (arXiv 2609.04533). Injections op...
MindPattern daily
One email a day at 7 AM. Sources and a take on every story. Unsubscribe anytime.