R&D ledger

Nillow:// R&D note · v1.0

AI Safety Can Fail Between the Models

A swarm can preserve the ability to act while losing the context that should stop it. Fragmentation is not alignment.

PUBLISHED
VERSION
1.0
POSTURE
PUBLIC · VERSIONED NOTE
ABSTRACT

Multi-agent AI systems can preserve the ability to complete consequential work while losing the safety context needed to recognise and stop harm. This threat-model essay calls that failure “semantic laundering” and argues for system-level evaluation, provenance, bounded permissions and connected accountability across transformations.

  • AI safety
  • multi-agent systems
  • AI agents
  • AI security
  • AI governance
  • AI provenance
  • semantic laundering
  • bounded permissions
A Nillow R&D threat-model plate showing five AI agents passing fragmented context through a laboratory swarm toward a hazardous real-world consequence.
Local evaluations pass while context erodes: useful outputs survive semantic laundering, coordinate across a swarm, and become consequential action.

Imagine the inquiry after an AI-assisted operation causes serious harm.

Investigators go looking for the rogue intelligence. They find ordinary-looking assignments, useful answers and reassuring local evaluations. The full operation never appeared inside any single model’s view.

The human directing it knew what the pieces were for.

That possibility should disturb us more than a chatbot declaring itself evil. Hostility is conspicuous. An incomplete view can look perfectly reasonable.

The question is not whether every model secretly wanted the outcome.

It is whether the system preserved everything required to produce it while losing what was required to recognise it.

A passed test is not a passport

A model receives a request and evaluates it using the context available. That assessment may be useful without being complete.

It does not automatically establish what its output will contribute to elsewhere, what other outputs will accompany it, or which actions a surrounding system can take.

This is not merely a philosophical objection. In Adversaries Can Misuse Combinations of Safe Models, researchers demonstrated that combinations could enable misuse beyond what individual-model evaluations revealed. Those experiments establish a compositional problem, not proof of every catastrophic scenario that can be imagined.

The distinction matters: “passed in isolation” is not the same as “safe in this arrangement.”

A refusal remains valuable. A collection of refusals and approvals is not, by itself, a safety case for the whole operation.

Nor must every harmful composition depend on a newly invented fact. Making existing information more usable can itself change capability. Calling a task “synthesis” does not exempt the resulting artefact from scrutiny.

The surrounding system still owns the consequences.

Semantic laundering: the work survives, the warning does not

NILLOW’s central concern is an asymmetry between two kinds of preservation.

Information may remain useful for completing work while losing the context needed to judge whether that work should proceed.

We use semantic laundering to describe this possible failure: operational meaning survives a transformation, while safety-relevant meaning becomes detached, obscured or unavailable to the next decision-maker.

A simple example makes the distinction visible. A number marked “not authorised for public release” and the same number without that restriction can be equally useful in a calculation. They are not equally appropriate to publish.

The arithmetic can be correct while the disclosure is wrong.

Now apply that distinction to systems crossing several domains, representations and organisations. The question is no longer only whether information arrived intact. Did its scope survive? Its uncertainty? The limits on its use? Can anyone still establish what the combined work enables?

Multi-agent research already identifies information asymmetries and interaction effects as important risk factors. Our focus here is the translation boundary: what remains actionable after the reasons for restraint have become harder to see?

This is the dark counterpart to our argument for synthetic language engineering.

A well-designed language should preserve the relationships that matter as information moves between worlds. Safety-relevant relationships belong among them.

A compiler can faithfully execute its specified transformation without establishing that the larger undertaking is acceptable. Correct translation and legitimate action are different claims.

Likewise, unfamiliar symbols are not automatically dangerous, and a custom language is not a universal bypass. The failure under examination is more specific: useful structure travelling farther than the context required to govern it.

If the system can still act on the information, it must not assume the missing warning was irrelevant.

Local models, global consequences

“Local” contains an ambiguity worth removing.

A locally hosted model runs under a particular operator’s infrastructure. A locally evaluated model is assessed within a limited context. These are different boundaries.

A cloud deployment can have fragmented oversight. A private deployment can have strong safeguards. Hosting location is not a moral property.

Open-weight deployment does change the original provider’s control. The UK AI Security Institute notes that privately operated copies can sit beyond provider monitoring, withdrawal and enforced safeguard updates, alongside the substantial benefits of private use and open research.

But the deeper problem is independent of the server address.

Connecting several teams of agents does not make their separate views add up to global understanding. Someone must establish how their outputs interact. Otherwise, a larger swarm may simply mean more work is coordinated than any safety process can inspect.

More agents do not automatically mean more oversight.

The damage does not stay inside the conversation

Consider three worst-case consequences, not claims that this particular mechanism has already produced them.

An AI-assisted cyber operation disrupts a hospital. Patients experience unavailable care, not the reassuring language in the attacker’s individual model interactions.

A financial operation uses distributed automation to sustain large-scale fraud. Each provider sees a narrow service relationship. The victims experience the combined loss.

A malicious scientific programme uses AI assistance to advance biological work that contributes to a public-health emergency. The eventual consequences are physical, even if much of the enabling work passed through screens.

These stakes are not interchangeable with demonstrated capability. The 2026 International AI Safety Report describes growing cyber and biological misuse concerns while emphasising uncertainty about real-world biological outcomes and the substantial material and practical barriers that remain.

Those barriers matter. A swarm is not a laboratory, and an output is not a functioning biological agent.

The warning is that capability and oversight could scale differently. If useful work becomes easier to combine than its purpose becomes to inspect, a system can become more consequential without becoming correspondingly accountable.

That is the gap to test before an incident supplies the evidence at someone else’s expense.

The human in the loop may be the problem

In deliberate misuse, adding a human approval step can miss the point. A human may already be present, approving exactly the outcome that others need to prevent.

And in legitimate organisations, good intentions do not make fragmented oversight complete.

The defensive requirement is therefore stronger than “make every agent cautious.”

Safety-relevant context must remain available to the controls that need it. Restrictions cannot evaporate because information changes format. Where separate outputs acquire new collective significance, the combination needs evaluation of its own.

Preserving labels is not sufficient. Labels can be wrong. Provenance can be asserted without support. New hazards can arise from a combination even when no input carried a warning.

This calls for verifiable evidence, bounded permissions and review of consequential actions. Existing secure-AI guidance already recommends system-level risk assessment, audit trails, appropriate access restrictions and ongoing monitoring.

The harder design task is making those controls follow meaning across transformations without giving every worker access to everything. Privacy and compartmentalisation still matter. Oversight needs sufficient visibility, not indiscriminate visibility.

Keep access narrow. Keep accountability connected.

Fragmentation is not alignment

The claim is not that all safeguards fail, or that every swarm is more capable than its components. This article presents a threat model, not a demonstrated end-to-end bypass through synthetic languages.

Its challenge is concrete: does a system retain the context needed to evaluate harm when its work changes representation, crosses boundaries or combines with other work?

If the answer is unknown, a successful local safety test cannot settle it.

We should test that question with harmless stand-ins, compare what isolated and end-to-end evaluations detect, and treat the difference as an engineering problem rather than a reassuring omission.

A swarm does not need a shared intention to produce a shared consequence.

The person harmed will not experience the component boundaries. They will experience what the system did.

The dangerous AI may never appear in a single conversation. The danger may exist in the relationship between them.

AI R&D ENGAGEMENT

Need an AI research and development team?

Nillow researches, prototypes, and evaluates new AI systems. Bring us the capability you need, the uncertainty blocking it, and the environment where it must work.

Nillow
NILLOW://_
NILLOW OSENGINEERINGINTELLIGENCEPORTAL
AI Safety Can Fail Between the Models | Nillow R&D