SemanTok exhibits high semantic alignment and video fidelity at every AR model size.
Generation vs. AR inference FLOPs per clip. Each faded curve is one AR size sweeping k from 1 to 256. Black: the best score each tokenizer reaches at a given compute. SemanTok's envelope is better over most of the compute range.
-
On Kinetics-600, SemanTok lowers gFVD by 11–24% and gFID by 5–17% and raises class accuracy by 25–61%, with the largest gains for the smallest AR models. On uCO3D, it improves gFVD, gFID and class accuracy by 2–13%, 5–8% and 22–30%.
-
This holds on both datasets, and in gFVD on Kinetics-600. The 85M SemanTok AR model matches that VideoFlexTok AR model in uCO3D gFVD and beats it in class accuracy on both datasets.
-
In class accuracy, ClipV and ViCLIP, on both datasets. On Kinetics-600, SemanTok's class accuracy is 0.631, against 0.560 for the 2.29B VideoFlexTok AR model.
-
More tokens cost more AR compute, and beyond a certain budget generations degrade in gFVD. VideoFlexTok's gFVD worsens after k=16, SemanTok's only after k=32.