Formula

= compliance with the content policy, = direct or indirect helpfulness.

The refusal paradigm classifies the prompt and fails on dual-use requests in both directions (o3 complied with a benign-sounding igniter question it refused when framed maliciously). Safe-completions judge the output: unsafe answers get 0, flat refusals get little, so the model learns safe, high-level or indirect help (used in GPT-5).

Appears in