Paper page - ImageWAM: Do World Action Models Really Need Video Generation, or Just Image Editing?
…In practice, ImageWAM does not decode the target frame at inference time; instead, it conditions a flow-matching action expert on the KV caches produced by image-editing denoising , using them as…