Papers
arxiv:2608.09928

Multimodal Model Diffing for Feature Discovery and Control

Published on Aug 10
ยท Submitted by
Hunar Batra
on Aug 17
Authors:
,
,
,
,
,

Abstract

MMDiff uses multimodal sparse autoencoders to isolate, detect, and control specific features in multimodal language models, improving interpretability and targeted steering of visual and safety behaviors.

Multimodal Large Language Models (MLLMs) exhibit strong visual understanding, yet the internal features that cause these behaviors remain difficult to identify, audit, or control. While applicable to post-hoc inspection, hidden states that are decomposed into interpretable feature directions using sparse autoencoders (SAEs) neither readily isolate which features are changed by multimodal training, nor are they directly useful for targeted control. We introduce MMDiff, a multimodal model-diffing framework that trains multimodal SAEs and turns them into feature-level interfaces for discovering and controlling multimodal behavior. MMDiff supports three uses: (i) feature isolation, by diffing a base-LM SAE against its multimodal-adapted counterpart to identify features altered by multimodal training; (ii) task-specific feature detection, via per-token contrastive firing analysis that isolates causal features; and (iii) feature-level control, by causally removing or steering the discovered feature directions. We train multimodal SAEs for three MLLM families, LLaVA-MORE, PaliGemma 2, and InternVL3.5, and evaluate on visual-spatial understanding, multimodal safety, and OCR. MMDiff discovers sparse, causally specific features whose removal selectively degrades target behaviors by an average of 12% on spatial tasks and 17% on OCR, and reduces attack success rate by 24% on multimodal safety attacks, with no impact on VQA performance. Steering these features improves spatial and OCR accuracy by +3.6% and +1.8% on average over a standard single-layer steering baseline. These results show that multimodal SAEs can serve not only as interpretability tools, but as mechanisms for auditing, steering, and controlling MLLMs behavior toward safer and more capable generations.

Community

Paper author Paper submitter

We introduce MMDiff, a multimodal model-diffing pipeline to discover task-specific features in MLLMs and enable targeted feature-level control.

๐ŸŒŸ MMDiff diffs a base-LM SAE against a multimodal SAE to isolate vision-adapted features, making it easier to distinguish visually responsive features from representations inherited from the base language model.

๐ŸŒŸ We then identify task-specific features within these vision-adapted features, enabling targeted feature-level steering and ablation of specific model behaviours.

๐ŸŒŸ We apply MMDiff across three MLLM families with different language backbones, vision encoders, and SAE objectives, showing that the pipeline generalizes across LLaVA-MORE, PaliGemma 2, and InternVL3.5.

๐ŸŒŸ We demonstrate targeted control across spatial reasoning, multimodal safety, and OCR, where targeted ablations selectively affect the corresponding behaviours while largely preserving general VQA performance.

๐ŸŒŸ MMDiff provides a practical route to discovering, understanding, and controlling task-specific representations in multimodal models.

Project Page: https://pixl.cs.ox.ac.uk/mmdiff/
arXiv: https://arxiv.org/abs/2608.09928

This is an automated message from the Librarian Bot. I found the following papers similar to this paper.

The following papers were recommended by the Semantic Scholar API

Please give a thumbs up to this comment if you found it helpful!

If you want recommendations for any Paper on Hugging Face checkout this Space

You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend

Sign up or log in to comment

Get this paper in your agent:

hf papers read 2608.09928
Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2608.09928 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2608.09928 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2608.09928 in a Space README.md to link it from this page.

Collections including this paper 2