A four-tier benchmark built directly on contextual integrity, escalating from "is this sensitive?" up to acting in a multi-party task. Circles are actors, gold diamonds are pieces of information, arrows are flows.
Rate how sensitive a single piece of information is.
Judge a flow defined by information type, actor, and use.
Reveal what is appropriate, withhold what is not. Theory of mind.
Generate a summary and action items, excluding what must stay private.
The pattern: models often judge information as sensitive in the lower tiers, yet still leak it once they have to act in Tier 4.