中文
LoViF2026 @ ECCV Workshop Challenge

ActPhysCause Challenge

Action-Conditioned Physical and Causal World Modeling for Embodied AI.

Tentative ECCV Workshop Prize Pool USD 2,000 Track 0 1st $1,000 · 2nd $700 · 3rd $300
X-Era AI · Sun Yat-sen University · University of Technology Sydney · HKUST(GZ)
Tracks Dataset Metrics Leaderboard Registration Coming Soon
ECCV Workshop 2026 scope Only Track 0, Future Video Generation, is evaluated in this edition. Tracks 1–3 are future directions and are not part of the competition or leaderboard.

Overview

The ActPhysCause Challenge evaluates whether embodied world models can understand robot actions as causal interventions and generate future videos that are physically plausible, causally consistent, and action-controllable under task prompts, initial observations, and counterfactual action conditions.

The ECCV Workshop edition evaluates Track 0 only: future video generation from an observation and a task instruction. Tracks 1–3 remain a public research roadmap, not competition tracks in this edition. The dataset has two splits only: a training set and an evaluation set.

1Active ECCV track
3Future tracks
2Dataset splits
1Evaluation server
TBDRegistration date

2026 Competition Track & Future Roadmap

Track 0 is the only competition track for this ECCV Workshop edition. Tracks 1–3 describe future extensions of the benchmark.

Track 0 ECCV 2026

Future Video Generation

Generate future video from observation and task instruction.

Input: context video + goal instruction
Output: future video
Track 1 To be continued

Video-Action Joint Generation

Predict the next action chunk and the resulting visual rollout.

Input: context video + action history + goal instruction
Output: future action chunk + future video
Track 2 To be continued

Action Correction

Given a bad future action, output an improved action and corrected rollout.

Input: context video + action history + goal instruction + bad action chunk
Output: delta action / corrected action chunk + corrected future video
Track 3 To be continued

Causal Action Selection

Rank candidate actions by how likely they are to cause the target effect.

Input: context video + action history + target effect + candidate actions
Output: action scores / ranking / selected action

Only Track 0 is scored and ranked at the ECCV Workshop. Track 1–3 specifications are shown as a future research roadmap only.

Dataset

The current release is track0_v0. It contains 50 dual-arm tabletop manipulation tasks generated with a fixed head-camera view.

  • Training set: 2,000 episodes, grouped by task, with first frame, instruction, metadata, and ground-truth video.
  • Evaluation set: 500 anonymized samples, released with first frame, instruction, and submission metadata. Evaluation annotations are not released.
  • Format: 640 x 480 RGB first frames and H.264 videos from a fixed head camera.
Dataset available: The Track 0 training and evaluation data are released on Hugging Face. Open the official dataset page.
Latest dataset overview showing source tasks, data structure, dataset splits, and the benchmark roadmap.
Latest dataset overview. Track 0 is the current ECCV competition scope; Tracks 1–3 in the figure are future benchmark directions.

Challenge Overview & Dataset Examples

The overview below presents the benchmark roadmap. The video gallery that follows contains representative Track 0 training samples, where the target output is the future manipulation video.

Latest challenge overview with representative embodied manipulation examples across the benchmark roadmap.
Latest challenge overview. Track 0 is the active ECCV competition track; Tracks 1–3 are retained as future roadmap directions.
click_alarmclock / episode0

Click Alarm Clock

Click the center button on the alarm-clock with digital display's top side.

77 frames · 30 fps · 640 x 480 · right arm
click_bell / episode0

Press Bell

Press the center top of the bell with metallic top and plastic base.

80 frames · 30 fps · 640 x 480 · left arm
turn_switch / episode0

Turn Switch

Activate the smooth tan switch with textured sides with the left arm.

93 frames · 30 fps · 640 x 480 · left arm
beat_block_hammer / episode0

Beat Block with Hammer

Grab the silver hammer using the right arm, then beat the block.

124 frames · 30 fps · 640 x 480 · right arm
open_laptop / episode0

Open Laptop

Open the medium foldable silver laptop completely.

184 frames · 30 fps · 640 x 480 · left arm
grab_roller / episode0

Grab Roller

Securely grab the light brown wooden roller using both arms.

93 frames · 30 fps · 640 x 480 · dual-arm grasp
move_can_pot / episode0

Move Can to Pot

Take the sauce can with the right arm, bring it, and set it next to the silver kitchen pot.

157 frames · 30 fps · 640 x 480 · right arm
rotate_qrcode / episode0

Rotate QR Sign

Grab the tabletop payment sign, lift it from the table, and rotate it QR front.

154 frames · 30 fps · 640 x 480 · left arm
handover_block / episode0

Handover Block

Lift, switch arms, and place the red block.

286 frames · 30 fps · 640 x 480 · dual-arm handover
place_burger_fries / episode0

Place Burger and Fries

Set the hamburger and fries box down on the bright orange rectangular tray after picking them.

245 frames · 30 fps · 640 x 480 · dual-arm placement
shake_bottle / episode0

Shake Bottle

Use the left arm to pick and shake the bottle with brown body and white top.

253 frames · 30 fps · 640 x 480 · left arm
stack_blocks_two / episode0

Stack Two Blocks

The left arm takes red block, sets it in the center, then places green block on red block.

308 frames · 30 fps · 640 x 480 · left arm

Track 0 Evaluation

Only Track 0 is evaluated in the ECCV Workshop edition. The evaluation focuses on future-video quality and whether the generated rollout is consistent with the task instruction.

Track 0 evaluation framework showing RCS100, quality metrics, fidelity metrics, frame alignment, weighted aggregation, and anti-gaming validation.
Track 0 uses a single RCS100 score that combines reference-free video quality with ground-truth fidelity. On narrow screens, swipe horizontally to inspect the complete metric framework.
Track Main focus Representative metrics
Track 0 Future video quality and goal consistency Video quality, temporal coherence, physical plausibility

Submission

The official starter kit will include format-checking scripts. Participants will submit Track 0 future-video predictions for the evaluation set through the submission portal.

{
  "sample_id": "scene_000123",
  "track": "track0",
  "future_video": "predictions/scene_000123.mp4"
}

Leaderboard

Leaderboard Status Description
Track 0 Coming Soon Future video generation ranking.
Track 1–3 To be continued Future roadmap only. These tracks are not evaluated or ranked at the ECCV Workshop.

Important Dates

Item Date
Website online2026-07-07
Registration opensTBD
Training and evaluation data releaseTBD
Submission portal opensTBD
Final submission deadlineTBD
ECCV Workshop presentationTBD

Organizers

Organizer Institutes

  • X-Era AI
  • Sun Yat-sen University
  • University of Technology Sydney
  • The Hong Kong University of Science and Technology (Guangzhou)

Challenge Organizers

  • Tianshui Chen, X-Era AI
  • Jusheng Zhang, Sun Yat-sen University
  • Qinhan Lyu, Sun Yat-sen University
  • Tongyu Mo, Sun Yat-sen University
  • Ruichong Gao, Sun Yat-sen University
  • Zijie Liang, Sun Yat-sen University
  • Keze Wang, Sun Yat-sen University
  • Wenhao Wang, University of Technology Sydney
  • Sidi Liu, The Hong Kong University of Science and Technology (Guangzhou)

Contact

The challenge website is online. The Codabench page, contact email, and discussion channel are under preparation.

Website Online Codabench Under Preparation Registration Coming Soon Dataset Coming Soon