One model · Nine tasks · Arbitrary camera control

OmniCamera

A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control

Yukun Wang1,2,* Ruihuang Li2,† Jiale Tao2 Shiyuan Yang2,3 Liyi Chen2,4 Zhantao Yang2 Handz2 Yulan Guo1,† Shuai Shao2 Qinglin Lu2

1 Sun Yat-sen University · 2 Hunyuan, Tencent · 3 CityU · 4 PolyU

* Work done during an internship at Tencent Hunyuan. Corresponding authors.

01 Camera control

Motion Text×Source Video
Dolly in

Re-camera an existing video with natural-language camera commands while preserving the source content.

02 Content input

The idea

Decouple what happens
from how it is seen.

Video intertwines two axes: the dynamic content of a scene and the camera motion through which it is observed. OmniCamera explicitly disentangles and commands both.

3 × 3control combinations
830K500K synthetic + 330K real videos
50predefined motion types
5Bparameter backbone

Method

Why one model handles all nine tasks.

Three ideas make the video control stable, precise, and photorealistic.

01

Unified architecture

Decoupled injection & Condition RoPE

Text prompts and visual conditions are fused through joint attention, while explicit camera trajectories are injected through a dedicated MLP pathway. 3D Condition RoPE assigns modality-specific spatial-temporal offsets to reduce interference.

02

OmniCAM dataset

Precision meets realism

Accurate UE5 trajectories provide geometric supervision; curated real videos restore photorealism and broaden scene diversity.

03

Training strategy

Dual-level curriculum

Training progresses from text to reference video to trajectory control, then from precise synthetic supervision to photorealistic real-world adaptation. Dual-condition classifier-free guidance provides separate scales for semantic text and camera-motion conditions.

Current scope. Finer-grained controls such as multiple reference images and localized motion guidance remain future work.

Interactive results

Nine tasks, explored in one place.

Each case exposes its camera condition, content condition, and generated result together. Videos load as they enter the viewport.

01 / 09

Text-controlled V2V

0 results

Citation

Build on OmniCamera.

If this work supports your research, please cite the paper.

View on arXiv
@misc{wang2026omnicamera,
  title         = {OmniCamera: A Unified Framework for Multi-task Video Generation with Arbitrary Camera Control},
  author        = {Wang, Yukun and Li, Ruihuang and Tao, Jiale and Yang, Shiyuan and Chen, Liyi and Yang, Zhantao and Handz and Guo, Yulan and Shao, Shuai and Lu, Qinglin},
  year          = {2026},
  eprint        = {2604.06010},
  archivePrefix = {arXiv},
  primaryClass  = {cs.CV},
  url           = {https://arxiv.org/abs/2604.06010}
}