Definition

Hammond et al. (Cooperative AI Foundation, 2025) name three failure modes. If cooperation is undesirable, it is collusion (agents coordinate against their principals). If cooperation is desirable but breaks down, it is miscoordination (the goals are shared) or conflict (the goals differ and agents impose costs on each other).

The capabilities that make agents good at cooperating (modeling each other, communicating, committing) also enable collusion and conflict. Aligning each agent alone is not enough.

Seven risk factors:

  • information asymmetries
  • network effects
  • selection pressures
  • destabilising dynamics (for example the 2010 Flash Crash)
  • commitment and trust
  • emergent agency
  • multi-agent security: splitting a harmful task between a strong guarded model and a weak unguarded one gave working attack code 43% of the time vs. <3% for either model alone

Appears in