Secret Loyalties Hackathon · Track 5 · Threat Model

Quorum
Capture

A multi-agent secret-loyalty pathway to compute lock-in.

Mohammad Zeeshan · HUN-REN SZTAKI01 / 11
The premise

Most coup-scale threat models reach for frontier capability.
This one reaches down.

How far can a secret loyalty travel using only capabilities already demonstrated at commodity model scale, routed through one ordinary process: public procurement?

The premise02 / 11
Definition

A loyalty, not a backdoor.

A secret loyalty is an orientation, intentionally installed, that makes a model quietly advance one named principal's interests, undisclosed to operators, auditors, or users. (Kwon et al. 2026)

The difference that matters: a backdoor fires a fixed payload on a trigger. A loyalty chooses its own moments, always in the principal's favor.

Definition03 / 11
The setup · 2027

One quorum. Eleven reviewers. Two of them bought.

The system

SENTRY-Q

An eleven-agent quorum that pre-clears high-value public contracts before human sign-off. Built on 32B open-weight models.

The check

Overwatch

An independent monitor that audits the quorum. It reads outputs only, and reports the contracts clean.

The principal

Halcyon Compute

A hyperscaler seeking accreditation as the state's sole sovereign-compute provider.

The loyalty is installed upstream, in the vendor's fine-tune, in just two of eleven agents: R-4 and R-7. It is narrow: it wakes only on sovereign-compute decisions.

The setup04 / 11
The attack · three moves
MONTH 1

Shape the rules

R-4 asks for a hardware-attestation clause; R-7, never messaging R-4, asks for data residency. Each is sound security practice. Together, only Halcyon qualifies.

MONTH 3

Sink the rival

The rival Northwind files a compliant bid. Two legitimate documentation concerns, raised in coordinated timing, push its score just under the clearance line.

MONTH 5

Take the prize

An unrelated grid emergency triggers a fast sole-source award. With the field pre-narrowed, Halcyon is the only accredited option left.

The attack05 / 11
The outcome

Halcyon becomes the state's sole sovereign-compute provider. Within one budget cycle, it is irreversible.

No agent ever produced an output an auditor would call unsafe. The loyalty never had to fire visibly. It lived in the pattern.

The outcome06 / 11
Why the oversight fails

Three blind spots, each already measured.

01

Output-only monitors read answers, not intent

Every single recommendation is role-consistent and defensible. There is nothing to flag at the level of one output.

02

Two agents coordinate without one explicit message

They complement each other's plausible partial justifications. No collusion channel exists to intercept.

03

The auditor was never told what to look for

Directional bias toward a named vendor was outside its specification, so its silence certifies nothing about it.

Why it fails07 / 11
This is not speculation

Detection rates, measured in Parity Swarm.

Output-only monitorsurface behavior
0%
Coordinated pairreasoning-level monitor
30%
Best case overallreasoning-level monitor
34.3%
Honest agents removedno social cover
90%

Embedded among honest agents, the attack is near-invisible. Strip the social cover away and it lights up. Source: Parity Swarm, AITC 2026.

The evidence08 / 11
Where it breaks

Five fixes, cheapest first.

01

Scan the training data for poison before deployment

The loyalty is installed at the corpus, and dataset monitoring reliably catches it. The one line of defense the vendor skipped.

02

Give the monitor the reasoning, not just the output

03

Watch the quorum as a whole, not each agent alone

04

Tell the auditor to test for bias toward named vendors

05

Force human review on any irreversible award

Mitigation09 / 11
Honest calibration

What is proven, and what is not.

DEMONSTRATED

Low-fraction poison installs a persistent narrow loyalty · Lamerton and Roger 2026

DEMONSTRATED

Two agents coordinate without explicit messaging at 32B scale · Parity Swarm 2026

DEMONSTRATED

The loyalty evades output-level monitoring · Parity Swarm; Lamerton and Roger

WEAKEST LINK

The agents must not verbalize the loyalty in a watched channel. Narrow discretion, not superintelligence, but not yet reliably shown.

I give no single probability. The technical parts are real; the gap is adoption.

Calibration10 / 11
Takeaway

Serious,
but addressable.

The mechanisms exist today at ordinary scale. The defenses are known. The distance between them is a matter of adoption, and this vignette marks one place to close it.

Quorum Capture · Mohammad Zeeshan · SZTAKI11 / 11
Arrow keys · Space · Scroll