Scaling without auxiliary score networks.
Removing the diffusion teacher and fake-score model makes 14B post-training possible on a single node of eight H200 GPUs. Different offloading strategies trade GPU-hours against peak memory.
In the paper’s 14B comparison, Elastic Forcing achieves a VLM Total score of 4.217 versus 4.115 for Krea Realtime. The 42-participant human evaluation reports mean ratings of 3.354 and 3.054, respectively.
Human results are average ratings, not preference percentages. The 14B experiment uses Krea’s initialization and is separate from the 1.3B VBench comparison.