MVHOI: Bridge Multi-View Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model

1The Australian National University 2CSIRO 3Baidu Inc. 4Sun Yat-sen University

*Equal contribution

MVHOI reenacts one source human-object interaction with different target objects.
Given only a few reference views of each target (left columns), our method preserves the source motion
while staying faithful to the target's geometry and appearance under rotation.

Abstract

Human–Object Interaction (HOI) video reenactment aims to transfer the interaction dynamics of a source video to a novel target object while preserving realistic hand–object coordination. Existing methods typically rely on sparse 2D motion controls and monocular references, which are insufficient for complex out-of-plane motion and large viewpoint changes. We present MVHOI, a two-stage framework combining implicit motion extraction, 3D-aware multi-view reasoning, and video generation. In the first stage, a motion extractor encodes object dynamics into implicit motion descriptors. Conditioned on these descriptors, our Motion-Driven Object Prior (MDOP) module queries a 3D foundation model over multi-view references of the target object and autoregressively predicts coarse object anchors — a sequence of images that track the object's evolving orientation and appearance under the source motion without any explicit pose estimation. In the second stage, a DiT-based video generation model uses these anchors as structural guidance and the multi-view references as appearance guidance. We further reuse cross-view attention from MDOP as a soft attention bias to reduce reference-view confusion. For long videos, a cross-iterative inference strategy refreshes subsequent object priors using refined video outputs. Experiments demonstrate consistent improvements over state-of-the-art methods in object fidelity, motion consistency, visual quality, and interaction realism.

Comparison with Baselines

Drag the divider to compare our result against GenHOI, the strongest existing baseline, or switch to VACE, HunyuanCustom or MimicMotion. All methods receive the same source video and the multi-view target references shown on the left; both sides play in sync.

Multi-view references of the target object
Target references
MVHOI (Ours) GenHOI

More Results

Source interactions reenacted with different target objects.

BibTeX

@article{tong2026mvhoi,
  title={MVHOI: Bridge Multi-view Condition to Complex Human-Object Interaction Video Reenactment via 3D Foundation Model},
  author={Tong, Jinguang and Wu, Jinbo and Wang, Kaisiyuan and Shen, Zhelun and Huang, Xuan and Xiang, Mochu and Li, Xuesong and Li, Yingying and Feng, Haocheng and Zhao, Chen and others},
  journal={arXiv preprint arXiv:2603.14686},
  year={2026}
}