Targeted Semantic Substitution Attack Breaks Vision-Language Model Robustness at Low Perturbation
New research demonstrates that Vision-Language Models (VLMs) can be manipulated using targeted semantic substitution attacks even at very low image or video perturbations ($\epsilon \leq 4/255$), undermining prior assumptions about model robustness in safety-critical scenarios.
Key Findings
- Attack Type
- Targeted semantic substitution in post-merger token space of Vision-Language Models
- Perturbation Range
- $\epsilon \leq 4/255$
- Image Impact
- 38% complete semantic replacement at $\epsilon = 4/255$
- Video Impact
- 35.9% complete semantic replacement at $\epsilon = 1/255$
- Phenomenon Observed
- Semantic fusion, where the LLM rationalizes conflicting visual signals
This research reveals that Vision-Language Models (VLMs), previously considered robust to adversarial attacks at low perturbations ($\epsilon \leq 4/255$), can in fact be successfully manipulated. The attack aligns the semantic representation of a source image with that of a target image inside the model, causing the VLM to misinterpret inputs even under strict evaluation criteria.

In images, semantic substitution can make the model recognize the target starting at $\epsilon = 2/255$; for $\epsilon = 4/255$ nearly 38% of attacks resulted in full replacement of the original content. In videos, the success rate for full semantic replacement reaches nearly 36% at a much lower perturbation threshold ($\epsilon = 1/255$).
The study further identifies 'semantic fusion' — a phenomenon where large language models rationalize and generate coherent text based on contradictory visual signals arising from such attacks.
