A multi-agent secret-loyalty pathway to compute lock-in.
How far can a secret loyalty travel using only capabilities already demonstrated at commodity model scale, routed through one ordinary process: public procurement?
A secret loyalty is an orientation, intentionally installed, that makes a model quietly advance one named principal's interests, undisclosed to operators, auditors, or users. (Kwon et al. 2026)
The difference that matters: a backdoor fires a fixed payload on a trigger. A loyalty chooses its own moments, always in the principal's favor.
An eleven-agent quorum that pre-clears high-value public contracts before human sign-off. Built on 32B open-weight models.
An independent monitor that audits the quorum. It reads outputs only, and reports the contracts clean.
A hyperscaler seeking accreditation as the state's sole sovereign-compute provider.
The loyalty is installed upstream, in the vendor's fine-tune, in just two of eleven agents: R-4 and R-7. It is narrow: it wakes only on sovereign-compute decisions.
R-4 asks for a hardware-attestation clause; R-7, never messaging R-4, asks for data residency. Each is sound security practice. Together, only Halcyon qualifies.
The rival Northwind files a compliant bid. Two legitimate documentation concerns, raised in coordinated timing, push its score just under the clearance line.
An unrelated grid emergency triggers a fast sole-source award. With the field pre-narrowed, Halcyon is the only accredited option left.
No agent ever produced an output an auditor would call unsafe. The loyalty never had to fire visibly. It lived in the pattern.
Every single recommendation is role-consistent and defensible. There is nothing to flag at the level of one output.
They complement each other's plausible partial justifications. No collusion channel exists to intercept.
Directional bias toward a named vendor was outside its specification, so its silence certifies nothing about it.
Embedded among honest agents, the attack is near-invisible. Strip the social cover away and it lights up. Source: Parity Swarm, AITC 2026.
The loyalty is installed at the corpus, and dataset monitoring reliably catches it. The one line of defense the vendor skipped.
Low-fraction poison installs a persistent narrow loyalty · Lamerton and Roger 2026
Two agents coordinate without explicit messaging at 32B scale · Parity Swarm 2026
The loyalty evades output-level monitoring · Parity Swarm; Lamerton and Roger
The agents must not verbalize the loyalty in a watched channel. Narrow discretion, not superintelligence, but not yet reliably shown.
I give no single probability. The technical parts are real; the gap is adoption.
The mechanisms exist today at ordinary scale. The defenses are known. The distance between them is a matter of adoption, and this vignette marks one place to close it.