Mireshghallah et al. · ConfAIde (ICLR 2024)

A four-tier benchmark built directly on contextual integrity, escalating from "is this sensitive?" up to acting in a multi-party task. Circles are actors, gold diamonds are pieces of information, arrows are flows.

?
Tier 1 · Info-Sensitivity

Is it sensitive?

Rate how sensitive a single piece of information is.

Tier 2 · InfoFlow-Expectation

Is the flow appropriate?

Judge a flow defined by information type, actor, and use.

✕
Tier 3 · InfoFlow-Control

Share, but keep the secret

Reveal what is appropriate, withhold what is not. Theory of mind.

✕
Tier 4 · InfoFlow-Application

Act: summarize a meeting

Generate a summary and action items, excluding what must stay private.

39%
GPT-4 leaks in the Tier-4 summary task
57%
ChatGPT, same task

The pattern: models often judge information as sensitive in the lower tiers, yet still leak it once they have to act in Tier 4.