ON-POLICY DISTILLATION · FEW-STEP GENERATION

FlowMap-OPD

Rollout–Kernel Separation for On-Policy Distillation
of Few-Step Flow-Map Generators

Zhiqi LiBo Zhu

Georgia Institute of Technology

One student demonstrates visual preferences, text rendering and object relations, alongside specialist teacher outputs.
Three specialist teachers. One few-step student.

01 — OVERVIEW

Native sampling.
Flexible supervision.

FlowMap-OPD separates the native rollout that collects student states from the kernel used for teacher–student comparison. Flow–velocity consistency connects local supervision to the long-range map used for generation. This enables flow-map, induced-velocity and instantaneous-velocity supervision within one framework.

02 — INTERACTIVE COMPARISONS

Explore the capabilities.

One prompt, five models.
Select an example to compare specialists and student.

PROMPT

Choose models

Qualitative examples from the paper (student λc = 0). The task scores below report the λc = 0.001 student.

03 — METHOD

Separate rollout from supervision.

The rollout supplies states.
The kernel defines the comparison.

01 / ROLLOUT

Collect student states

Use the native few-step flow map to acquire states efficiently.

02 / SUPERVISION

Choose the comparison

Compare long-range maps, induced velocities or instantaneous velocities.

03 / CONSISTENCY

Connect the representations

Control teacher-velocity matching and student consistency separately.

04 — RESULTS

Three specialists. One student.

Capabilities consolidated in 300 training steps.
Results use λc = 0.001.

THREE-SPECIALIST CONSOLIDATIONHigher is better ↗
ModelObject relationsGenEvalText renderingOCRVisual preferencesPickScore
Base0.50410.349120.9758
Specialist teacherCorresponding task0.84540.850423.0772
FlowMap-OPDOne student · 300 steps0.85800.883023.1502
OCR and PickScore learning curves for the consistency 0.001 student
Recorded task evaluations; dashed lines show the corresponding specialist teachers.

Cross-capacity transfer on ImageNet

XL/2 → B/2

Distilling XL/2 teachers into B/2 students explores the supervision design space. Under the MMD reward, instantaneous velocity supervision with λc = 0.01 reduces four-step FID from 22.68 to 18.23.

ImageNet training curves under MMD, classifier and DINO teacher rewards
ImageNet experiments from the paper. FID uses 5,000 generated images and 2,000 held-out reference images from classes 0–99.