PROWBench tests whether generated videos show the actions and end states prescribed by a program, using replayable world records as the reference.
- 2026/10/11 — Released the dataset with Viewer image previews and the representation-swap tool with runnable scripts and a verified before-and-after example.
- 2026/10/05 — The paper was updated on arXiv.
Run these commands from the repository root:
# 1. Install (Python 3.11)
pip install -U huggingface_hub
pip install -r eval/requirements.txt
# 2. Download
hf download AlayaLab/PROWBench --repo-type dataset --local-dir PROWBench
# 3. Prepare
python3 tools/prepare_eval_data.py --download PROWBench --out prowbench-eval-dataprowbench-eval-data/ uses symbolic links to the download; keep PROWBench/ at its original location.
PROWBench's core data are engine recordings of scene states, object motion and camera trajectories, which serve as the
reference for evaluation. Use these records to create or adapt representations and write prompts suited to your model's
input requirements. We provide example representation videos (white, mixed and colored-obb) and prompts as starting points.
See the representation-swap tool for runnable code, a reusable skill and a before-and-after demo of changing visible geometry while preserving the recorded camera and object trajectories.
1. Generate videos. Generate one video per view using your prepared inputs and, when needed, an available first-frame reference. Evaluate one representation at a time and save outputs using this path (example below):
model_output_video/<dataset>/<method>/<episode>/<camera>/<representation>/video.mp4
model_output_video/verified_ff_40/my_model/dv200-001-city-crossing/follow/mixed/video.mp4
Use the episode ID, camera and representation from prowbench-eval-data/<dataset>/samples.jsonl.
Video requirements.
2. Evaluate. Set up paths and environments, then follow your task guide:
| Dataset | Guide |
|---|---|
verified_ff_40, unverified_ff_90, gunfight |
Five-second tasks |
long_horizon |
Long video (30 s) |
multi4 |
Multi-view (5 s) |
3. Collect results. Final tables are written to $WORK/table.json; multi-view appearance and geometry scores
are in $WORK/mv/summary.json. CPU example · Resuming runs.
| Dimension | Metrics |
|---|---|
| Entity control | IoU, CErr, Miss |
| Camera control | Rot, Trans, CamMC |
| Logic and state | ISR, LRA, gISR, gLRA |
| Visual and temporal quality | Img, Aes, Smooth, Warp, CLIP |
| Multi-view consistency | Comp., Cons., MEt3R |
| Long-horizon memory | Re-IoU, Re-State |
Definitions, units and aggregation rules: evaluation reference.
- Paper and project page
- Engine recordings, example representation videos and prompts
- Evaluation code and usage guide
- Increase dataset diversity
- Release the World Generation Pipeline
The code and dataset are released under the MIT License.
If you find PROWBench helpful for your research, please consider citing our paper:
@misc{huang2026prowbench,
title={PROWBench: Do Video Models Render What the Program Specifies?},
author={Zheng-Hui Huang and Guixu Lin and Yu-Ju Tsai and Jian-Kai Zhu and Fengbo Lan and Yu-Lun Liu and Yung-Yu Chuang and Kaipeng Zhang and Zhixiang Wang},
year={2026},
eprint={2610.02205},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2610.02205}
}