Learning the Target PriorsBefore Image Translation

A Decoupled Training Paradigm for Cross-Modal Image Translation in Remote Sensing

1Central South University2Wuhan University3Guangzhou University
* Equal contribution Corresponding authors
LTP-BIT: learning target priors before image translation

Our experiments reveal a systematic link between prior generation and downstream translation: under a fixed adaptation protocol, improvements in prior-generation FID are mirrored by lower translation FID as model and data scale increase.

LTP-BIT turns this finding into a prior-first framework that freezes the pretrained DiT and learns lightweight, task-specific cross-modal control through P-DART. With only 9.81% task-specific parameters, it achieves state-of-the-art overall performance across three SAR-to-RGB and NIR-to-RGB benchmarks, while retaining near-full-data instance fidelity on QXS-SAROPT with only 25% of the paired samples. The visual results show that learning the target-domain prior before translation enables P-DART to guide source-conditioned generation toward more realistic target images.

QXS-SAROPT

Single-polarization SARRGB

SpaceNet6

Multi-polarization SARRGB

Chesapeake

NIRRGB

Citation