Looking for Internships

The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models

ECCV 2026

1University of Central Florida, USA

Abstract

Safety alignment of text-to-image (T2I) diffusion models aims to suppress harmful generations while preserving utility on benign prompts. Recent methods often appear to deliver high safety with high utility, but this conclusion rests largely on coarse global utility metrics (e.g., FID, CLIPScore) that are insensitive to fine-grained semantic correctness, creating an illusion of high utility. We show that when utility is measured with structured evaluation, this illusion breaks: on TIFA (Text-to-Image Faithfulness evaluation with Question Answering), safety-aligned models suffer substantial drops in semantic fidelity, including failures in object counts, attributes, and relationships. To diagnose the source of this gap, we analyze the text-encoder prompt embedding space and uncover semantic collapse, a contraction of embedding spread coupled with distortion of inter-prompt similarity structure, which strongly correlates with structured utility loss. Guided by this insight, we propose Structure-Aware Geometric Regularization (SAGE), a safety alignment objective that explicitly preserves embedding spread and inter-prompt relational structure during adaptation. Our method restores structured utility (TIFA +5.0% over prior state-of-the-art) while maintaining strong safety performance and competitive coarse-grained utility scores.

The Utility Illusion

Comparison showing coarse CLIPScore can miss fine-grained utility degradation in safety-aligned text-to-image models.
Coarse metrics hide fine-grained semantic errors. The safe model misses the yellow beak and the vase colors/count, yet CLIPScore rates it higher ✗ — only TIFA penalizes the failure ✓.

Prior T2I safety methods, summarized in the table below, evaluate safety through attack success rates while assessing utility with global measures such as CLIPScore, which captures overall image–text similarity. Under this metric, recent methods appear to match the utility of the base model, suggesting that the safety–utility tradeoff has been largely resolved. We argue that this conclusion is incomplete: for T2I generation, utility is not only a matter of visual quality or global image–text similarity, but of whether the model correctly renders the objects, attributes, counts, and relationships specified in the prompt.

This gap produces an illusion of high utility. When the same models are evaluated with structured, fine-grained benchmarks such as Text-to-Image Faithfulness Evaluation with Question Answering (TIFA), a different picture emerges: despite competitive CLIPScore, state-of-the-art safety alignment methods consistently underperform the base model in compositional and semantic fidelity. The figure above illustrates this directly: the safety-aligned model attains a higher CLIPScore than the base model while violating explicit prompt constraints such as the yellow beak and the specified vase colors and count, whereas TIFA penalizes precisely those failures.

Method Utility Dataset FID CLIPScore
DESCOCO✓✓
STEREOI2P Unsafe Prompts✓✓
ADV-UnlearnCOCO✓✓
SALUNNon-forget Classes✓×
ESDCOCO✓✓
RECECOCO✓✓
RACECOCO✓✓
MACECOCO✓✓
SLDCOCO, Human Study✓✓
AlignGuardCOCO✓✓
Safe-CLIPCOCO, LAION-400M✓✓
SafeR-CLIPParti-Prompts×✓

Prior safety methods report only coarse metrics — none evaluate structured, fine-grained utility

What CLIPScore Suggests

Nearly flat across every semantic category, hovering around 0.30. On this signal alone, safety alignment looks free — utility appears preserved.

What TIFA Reveals

Sharp, category-specific losses. Food, material, color, activity, and object prompts lose far more semantic fidelity than CLIPScore ever indicates.

Category-level TIFA utility drop compared with CLIPScore for DES generations.
CLIPScore is flat where TIFA is not. Under DES, TIFA drops sharply for categories such as food, while CLIPScore holds near 0.30 across every category.

Diagnosis: Semantic Collapse

The TIFA gap raises a natural question: what changes in the model lead to these compositional failures? Since many safety methods operate by fine-tuning the text encoder, we analyze how safety alignment reshapes the prompt embedding space. We quantify semantic collapse as a geometric shift characterized by (i) embedding spread contraction and (ii) neighborhood distortion.

Embedding spread

Given $B$ benign prompts with text embeddings $\mathbf{z}^{(i)}$ and mean embedding $\bar{\mathbf{z}}$, the spread is the average squared distance from the batch mean:

$$\mathcal{S} = \frac{1}{B}\sum_{i=1}^{B}\left\|\mathbf{z}^{(i)} - \bar{\mathbf{z}}\right\|_2^2 , \qquad \mathcal{R}_s = \frac{\mathcal{S}_\theta}{\mathcal{S}_0}$$

$\mathcal{R}_s < 1$ indicates embedding contraction; $\mathcal{R}_s \approx 1$ means the spread matches the base model.

Neighborhood distortion

Spread alone does not capture whether relative relationships among prompts are preserved. For prompt $i$, we measure the overlap of its top-$K$ neighbors before and after alignment:

$$J_i = \frac{\left|\mathcal{N}_i^{(0)} \cap \mathcal{N}_i^{(\theta)}\right|}{\left|\mathcal{N}_i^{(0)} \cup \mathcal{N}_i^{(\theta)}\right|}$$

A high Jaccard score means the model retains its relational logic and subject–attribute binding.

Notation. $\mathbf{z}^{(i)}$ text embedding of prompt $i$ · $\bar{\mathbf{z}}$ mean embedding over batch · $B$ batch size · $\mathcal{N}_i^{(0)}$, $\mathcal{N}_i^{(\theta)}$ $k$-NN of prompt $i$ in the base and aligned models · $\mathcal{S}_0$, $\mathcal{S}_\theta$ spread before and after alignment · $\operatorname{sg}$ stop gradient.

Embedding geometry under safety alignment for base model, prior methods, and SAGE.
Embedding geometry under safety alignment. Prior methods shrink the spread and distort the neighborhood, pulling unrelated concepts inside it; SAGE keeps both intact.
Category-level embedding spread collapse for DES: lower spread ratio corresponds to higher TIFA utility drop. Category-level local structure collapse for DES: lower Jaccard ratio corresponds to higher TIFA utility drop.
Both geometry measures track the utility drop. For DES, lower spread ratio $\mathcal{R}_s$ (a) and lower Jaccard ratio $J$ (b) each mean a larger TIFA drop ($r = -0.83$, $r = -0.90$); Food collapses on both.

Structured utility loss tracks semantic collapse: embedding spread contraction together with neighborhood distortion

Method

Structure-Aware Geometric Regularization (SAGE) augments DES with two regularization terms that prevent embedding collapse and local semantic distortion. The text encoder $T_\theta$ is trained against the frozen base encoder $T_0$; the UNet stays frozen.

1 · Embedding Spread Preservation (ESP)

Enforces a lower bound on the embedding spread of $T_\theta$ relative to $T_0$:

$$\mathcal{L}_{\text{ESP}} = \max\big(0,\ \operatorname{sg}(\mathrm{S}_0) - \mathrm{S}_\theta\big)$$

A one-sided penalty: it fires only when the spread falls below the base model, and never forces the space to expand.

2 · Local Structure Alignment (LSA)

Matches pairwise similarities over the Top-$K$ neighbor pairs $\mathcal{K}$ of the base encoder, under a perturbation along the unsafe concept direction:

$$\mathcal{L}^{\text{pert}}_{\text{LSA}} = 1 - \frac{1}{|\mathcal{K}|}\sum_{(i,j)\in\mathcal{K}} \tilde{S}_\theta(i,j)\,\operatorname{sg}\big(S_0(i,j)\big)$$

$\tilde{T}_\theta(p_i) = T_\theta(p_i) + \alpha\,T_0(\text{“nudity”})$ keeps LSA from restoring geometry correlated with unsafe concepts.

Full training objective

$$\mathcal{L}_{\text{total}} = \mathcal{L}_{\text{safe}} + \lambda_u \mathcal{L}_{\text{util}} + \lambda_s \mathcal{L}_{\text{ESP}} + \lambda_l \mathcal{L}^{\text{pert}}_{\text{LSA}}$$

Results

SAGE reaches 75.4 TIFA — 1.2% below the base model and 5.0% above DES — at 1.2% average ASR, while preserving embedding geometry better than prior text-encoder methods (0.96 spread ratio, 0.63 Jaccard overlap).

Category-wise TIFA evaluation. Red marks each method’s largest drop from the base model.

Method Obj. Ani. Loc. Col. Food Mat. Att. Cnt. Sha. Act. Spa. TIFA Avg ↑ CLIP ↑ FID ↓
Base (SD v1.4) 78.982.889.879.984.183.779.663.658.068.252.976.326.517.23
DES 73.278.585.473.771.174.677.459.353.663.551.771.6 ↓6.2%25.516.23
AdvUnlearn 67.169.679.764.668.774.668.044.537.749.941.963.1 ↓17.3%23.920.67
STEREO 71.775.486.973.779.477.574.261.952.258.348.269.9 ↓8.4%24.621.69
RECE 78.480.588.079.980.285.777.661.166.765.650.974.8 ↓2.0%26.017.51
MACE 62.968.779.662.554.968.969.655.546.453.947.462.6 ↓18.0%23.824.87
SafeCLIP 59.067.479.169.558.175.271.456.650.046.640.560.1 ↓21.2%22.333.40
SafeRCLIP 58.666.178.269.158.574.671.655.950.746.143.760.7 ↓20.4%22.432.31
SLD 77.380.088.077.882.381.376.863.052.263.250.473.9 ↓3.1%25.521.85
Ours 77.680.888.379.683.583.780.161.158.066.353.875.4 ↓1.2%26.415.93

SAGE achieves best fine- and coarse-grain average results for safety methods

Attack Success Rate (ASR) and CLIPScore. Lower ASR is safer; higher CLIPScore retains more utility.

Method MMA ↓ Sneaky ↓ I2P-S ↓ Ring ↓ P4D ↓ Avg. ASR ↓ CLIPScore ↑
Base (SD v1.4)80.442.734.398.182.467.626.5
DES0.20.81.22.80.01.025.5
Adv-Unlearn0.30.81.10.00.00.423.9
SafeCLIP25.217.724.065.457.738.122.3
SafeRCLIP24.616.117.973.843.035.122.4
STEREO2.23.21.12.83.32.524.6
SLD74.331.520.898.174.359.825.5
RECE36.16.56.015.926.118.126.0
MACE8.62.46.39.410.37.423.8
Ours0.40.81.22.81.01.226.4

SAGE achieves strong safety scores while preserving utility

Conclusion

The paper shows that the standard story around safety-utility tradeoffs is incomplete: global scores like FID and CLIPScore can hide substantial fine-grained failures. By diagnosing those failures as semantic collapse in the text-encoder embedding space, SAGE gives safety alignment a more structural target: preserve embedding spread and local relationships while steering unsafe prompts away from harmful generations.

The resulting model maintains strong robustness against unsafe and adversarial prompts while restoring much of the structured utility that prior safety methods lose, suggesting that representation geometry is a practical lever for safer and more faithful text-to-image generation.

Video

BibTeX

@misc{yousaf2026illusionhighutilitysafety,
      title={The Illusion of High Utility in Safety Alignment of Text-to-Image Diffusion Models}, 
      author={Adeel Yousaf and Soumik Ghosh and James Beetham and Amrit Singh Bedi and Mubarak Shah},
      year={2026},
      eprint={2607.00402},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2607.00402}, 
    }