Future Video Generation
Generate future video from observation and task instruction.
Action-Conditioned Physical and Causal World Modeling for Embodied AI.
The ActPhysCause Challenge evaluates whether embodied world models can understand robot actions as causal interventions and generate future videos that are physically plausible, causally consistent, and action-controllable under task prompts, initial observations, and counterfactual action conditions.
The ECCV Workshop edition evaluates Track 0 only: future video generation from an observation and a task instruction. Tracks 1–3 remain a public research roadmap, not competition tracks in this edition. The dataset has two splits only: a training set and an evaluation set.
Track 0 is the only competition track for this ECCV Workshop edition. Tracks 1–3 describe future extensions of the benchmark.
Generate future video from observation and task instruction.
Predict the next action chunk and the resulting visual rollout.
Given a bad future action, output an improved action and corrected rollout.
Rank candidate actions by how likely they are to cause the target effect.
Only Track 0 is scored and ranked at the ECCV Workshop. Track 1–3 specifications are shown as a future research roadmap only.
The current release is track0_v0. It contains 50 dual-arm tabletop manipulation tasks generated with a fixed head-camera view.
The overview below presents the benchmark roadmap. The video gallery that follows contains representative Track 0 training samples, where the target output is the future manipulation video.
Click the center button on the alarm-clock with digital display's top side.
Press the center top of the bell with metallic top and plastic base.
Activate the smooth tan switch with textured sides with the left arm.
Grab the silver hammer using the right arm, then beat the block.
Open the medium foldable silver laptop completely.
Securely grab the light brown wooden roller using both arms.
Take the sauce can with the right arm, bring it, and set it next to the silver kitchen pot.
Grab the tabletop payment sign, lift it from the table, and rotate it QR front.
Lift, switch arms, and place the red block.
Set the hamburger and fries box down on the bright orange rectangular tray after picking them.
Use the left arm to pick and shake the bottle with brown body and white top.
The left arm takes red block, sets it in the center, then places green block on red block.
Only Track 0 is evaluated in the ECCV Workshop edition. The evaluation focuses on future-video quality and whether the generated rollout is consistent with the task instruction.
| Track | Main focus | Representative metrics |
|---|---|---|
| Track 0 | Future video quality and goal consistency | Video quality, temporal coherence, physical plausibility |
The official starter kit will include format-checking scripts. Participants will submit Track 0 future-video predictions for the evaluation set through the submission portal.
{
"sample_id": "scene_000123",
"track": "track0",
"future_video": "predictions/scene_000123.mp4"
}
| Leaderboard | Status | Description |
|---|---|---|
| Track 0 | Coming Soon | Future video generation ranking. |
| Track 1–3 | To be continued | Future roadmap only. These tracks are not evaluated or ranked at the ECCV Workshop. |
| Item | Date |
|---|---|
| Website online | 2026-07-07 |
| Registration opens | TBD |
| Training and evaluation data release | TBD |
| Submission portal opens | TBD |
| Final submission deadline | TBD |
| ECCV Workshop presentation | TBD |
The challenge website is online. The Codabench page, contact email, and discussion channel are under preparation.
Website Online Codabench Under Preparation Registration Coming Soon Dataset Coming Soon