Paper page - Vision as Unified Multimodal Generation
… The following papers were recommended by the Semantic Scholar API UniVL: Unified Vision-Language Embedding for Spatially Grounded Contextual Image Generation 2026 UniCanvas: A Diffusion-base Unified Model for Text-in-Image Joint Generation 2026 An Open-Source Benchmark and Baseline for Multi-tempor… …