SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation

Mikhail Dereviannykh1,2 Vikram Voleti1 Simon Donné1 Mallikarjun Byrasandra Ramalinga Reddy1 Shimon Vainer1 Mark Boss1
1Stability AI 2Karlsruhe Institute of Technology
Paper (soon) Code (soon) Videos 7 Findings

Text-to-video · uCO3D

Prompt: “A small orange basketball on a plaid tablecloth”

At k=4, the 201M SemanTok AR model already keeps the ball's shape and appearance through the orbit. VideoFlexTok's ball is misaligned at the same size, and still unstable up to k=64 with an 11× larger AR model.

Class-to-video · Kinetics-600

Class: “yoga” · both AR models 2.29B (equal cost)

SemanTok keeps a complex body motion stable from k=16; larger k refines it. VideoFlexTok changes the scene between k=4 and k=16.

TL;DR

Flexible video tokenizers (e.g., VideoFlexTok) let an autoregressive (AR) model stop after any number of tokens, which condition a diffusion decoder, so the first tokens should already capture what the clip shows. SemanTok supervises this explicitly: every nested token prefix is trained to carry the clip's semantics. The resulting prefixes are cheaper to predict and lead to better generation fidelity and higher semantic alignment: a 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size.

3.4× smallera 201M SemanTok AR model matches or beats a VideoFlexTok AR model 3.4× its size
+25–61%semantic alignment on Kinetics-600 vs. the same-size VideoFlexTok AR model
+11–24%fidelity on Kinetics-600 vs. the same-size VideoFlexTok AR model
All metricsimprove at high k: at k=256, fidelity and semantic alignment beat VideoFlexTok on both datasets at every AR size

k = number of tokens per latent frame that the AR model generates (its token budget), ordered coarse to fine: .

How it works

Previous flexible tokenizers (e.g., VideoFlexTok) encode each latent frame into 256 ordered tokens, and any prefix of k tokens per frame decodes to a video. The AR model predicts tokens coarse to fine, up to the budget k chosen at inference.

We introduce two blocks on top of it:

Semantic EncoderFrozen DINOv2 patch features are concatenated with each VAE latent patch before the input projection. A zero-initialized projection of each frame's DINO class token is added to that frame's first register token.
Semantic AlignmentA Dense and a Class DINO head, each two cross-attention layers, reconstruct the DINO patch features and class token from the kept token prefix alone. Unlike decoder REPA, they never see the noised latent, so only the tokens can lower their loss.

Everything else is unchanged: SemanTok keeps VideoFlexTok's FSQ codebook (64k codes), sequence length, nested dropout, native decoder and its losses, so both tokenizers are compared under the same AR models and training recipe.

No loss assigns information to particular tokens; nested dropout naturally concentrates the most relevant information in the earliest ones.

Full SemanTok tokenizer

Orange: the VideoFlexTok path. A frozen VidTok VAE maps the clip to latents; a time-causal encoder reads patches and K=256 learnable register tokens per frame; FSQ quantizes the register outputs; nested dropout keeps a token prefix that conditions a time-causal rectified-flow decoder, trained with flow matching and decoder REPA. Purple: SemanTok's semantic supervision. DINOv2 patch features are concatenated with each VAE patch, the frame's DINO class token is added to register token rt,0, and the Dense and Class DINO heads cross-attend to the kept prefix of frames ≤ t (cosine losses, weight 0.5 each).

Seven findings

We compare the two tokenizers at seven AR sizes (49M–2.29B) and token budgets k=1–256, on class-to-video (Kinetics-600) and text-to-video (uCO3D), for fidelity (gFVD, gFID) and semantic alignment (class accuracy, ViCLIP, ClipV).

Click a ● point to see the numbers behind it.

1

SemanTok exhibits high semantic alignment and video fidelity at every AR model size.

Generation vs. AR inference FLOPs per clip. Each faded curve is one AR size sweeping k from 1 to 256. Black: the best score each tokenizer reaches at a given compute. SemanTok's envelope is better over most of the compute range.

  • On Kinetics-600, SemanTok lowers gFVD by 11–24% and gFID by 5–17% and raises class accuracy by 25–61%, with the largest gains for the smallest AR models. On uCO3D, it improves gFVD, gFID and class accuracy by 2–13%, 5–8% and 22–30%.
  • This holds on both datasets, and in gFVD on Kinetics-600. The 85M SemanTok AR model matches that VideoFlexTok AR model in uCO3D gFVD and beats it in class accuracy on both datasets.
  • In class accuracy, ClipV and ViCLIP, on both datasets. On Kinetics-600, SemanTok's class accuracy is 0.631, against 0.560 for the 2.29B VideoFlexTok AR model.
  • More tokens cost more AR compute, and beyond a certain budget generations degrade in gFVD. VideoFlexTok's gFVD worsens after k=16, SemanTok's only after k=32.
2

Increasing SemanTok model size further improves fidelity.

Best-k scores vs. AR model size. Fidelity scales with AR size for both tokenizers, while the semantic-alignment gap between them does not close.

  • From 49M to 2.29B, SemanTok's best-k gFVD falls from 224 to 202 on Kinetics-600 and from 218 to 197 on uCO3D. Larger AR models mostly trim the compounding error of the token rollout.
  • SemanTok's best-k class accuracy and ClipV lie above VideoFlexTok's at every size. SemanTok's benefit is largest where AR capacity is scarce: at k=64, its gFID gain over VideoFlexTok halves from 49M to 2.29B.
  • For the 1.33B AR model on Kinetics-600, SemanTok's best-k gFVD lead shrinks from 28% to 9% by 26B training tokens, then holds at 11–14%. Its class-accuracy lead is still 30% after 65.5B tokens. On uCO3D, SemanTok's ClipV nearly saturates by 13B tokens and stays above VideoFlexTok's.

One rollout per tokenizer · class-to-video, “playing guitar”, by AR model size

3

SemanTok is able to maintain semantic alignment over out-of-distribution classes.

Flashlight · in-distribution class (ID)

Fedora · out-of-distribution class (OOD)

uCO3D reconstructions from ground-truth tokens. SemanTok recovers the object class at a smaller k, on a seen (ID) and an unseen (OOD) class.

  • OOD object classes are never seen in training. At k=16 on uCO3D, SemanTok's reconstructions score higher ClipV and class accuracy than VideoFlexTok's, but about 2.7 dB lower PSNR.

    Reconstruction from ground-truth tokens on uCO3D, VideoFlexTok / SemanTok; relative gain below, better score in bold. Class accuracy is NCM. 152 OOD and 1,014 ID validation clips; tokenizers at 100k steps. Beyond k=64–128, VideoFlexTok's better-reconstructing tokens catch up.

  • With a 201M AR model at k=16, SemanTok raises generated class accuracy over VideoFlexTok by 24% on ID clips and by 29% on OOD classes. The same holds at every AR size: from k=2 to 256, SemanTok's class accuracy is higher by 10–36% on OOD classes and by 9–31% on ID clips (at k=1 the two are within −6% to +12%). At 201M, ViCLIP favours SemanTok from k=16 on, but not at k=4.

    Generated class accuracy on uCO3D, VideoFlexTok / SemanTok; relative gain below. 2,560 generated clips per split and cell; AR step 20k, CFG 3.0, seed 0.

4

SemanTok achieves higher decoder-REPA semantic alignment at all noise levels, including the pure noise setting.

Readout cosine to DINOv2 vs. k (left) and latent gain, σ=0.25 minus σ=1 (right).

VideoFlexTokvsSemanTok
input
latent input
DINOv2 target
0
  • Decoder REPA aligns an early decoder layer with DINOv2 features, but that layer also sees a partly noised latent, which can supply part of the target without the tokens. SemanTok keeps decoder REPA and adds the DINO heads, which read only the tokens.
  • At k=32 on uCO3D, SemanTok's pure-noise readout reaches the similarity that VideoFlexTok reaches only with a 75%-clean latent (0.721 vs. 0.716, circled).
  • At k=256 the latent adds almost nothing to SemanTok's readout, but over 11× more to VideoFlexTok's. The exception is the smallest budgets with a mostly clean latent (Finding 7).
5

SemanTok has high fidelity on reconstruction as well as generation.

  • The realization gap is what a tokenizer loses from reconstruction (ground-truth tokens) to generation (AR-sampled tokens).
  • Larger AR models narrow the gap between generation and reconstruction for both tokenizers, but none closes it.
  • On Kinetics-600, the AR model sees the class label; a reconstruction does not.
6

SemanTok's generation fidelity gain comes from its earlier tokens, which are cheaper to predict.

(a) Teacher forcing · first m tokens per frame forced

tokens of one frame, coarse → fine (1 … 256)
m=0
m=16
m=64
ground-truth (forced)AR-sampled

(b) Cross-entropy at position · token t of a frame

tokens of one frame, coarse → fine (1 … 256)
t=3
t=6
ground-truth context?token being predicted

CE(t) = −log2 p(zt | z<t, class)

(c) Prefix cost · first k tokens per frame

tokens of one frame, coarse → fine (1 … 256)
k=3
k=6
prefix: averagednot counted

prefix CE(k) = 1⁄k Σt≤k CE(t)

  • With nothing forced, SemanTok's gFVD is 23% lower than VideoFlexTok's at 201M. Once the first 16–64 tokens are forced, the lead vanishes, and VideoFlexTok's better-reconstructing tail edges ahead.
  • Cross-entropy of 8.8 vs. 12.9 bits per token at 201M. SemanTok's marginal entropy is only about one bit lower, so most of the saving comes from context. It does not come from repetition: SemanTok repeats tokens less often overall than VideoFlexTok.
  • SemanTok's first 4 tokens cost 37 bits per frame, yet nearly match the class accuracy of VideoFlexTok's first 32 tokens (405 bits). Later tokens restore part of that detail, and the generative decoder fills in the rest.
7

SemanTok with one token per frame struggles to serve every objective.

SemanTok's advantage over VideoFlexTok at equal AR size vs. k, signed so that positive values favour SemanTok. Bands: 95% paired-bootstrap intervals over evaluation clips; the grey line marks no difference.

  • On Kinetics-600, SemanTok is never significantly worse. SemanTok's fidelity lead follows at larger budgets (Finding 1).
  • At k=1, SemanTok's pure-noise readout leads by only 0.01–0.02 DINOv2 cosine, against 0.04–0.08 from k=16. With a mostly clean latent (σ=0.25), SemanTok's readout is lower up to k=4 on uCO3D and k=8 on Kinetics-600.
  • SemanTok still never falls behind in class accuracy, at any budget or AR size.

Limitation. Because SemanTok prioritizes semantic alignment over reconstruction, it struggles to reconstruct the same colours and appearance details at lower token budgets.

More videos

Click a video for controls.

BibTeX

@misc{dereviannykh2026semantok,
  title  = {SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation},
  author = {Dereviannykh, Mikhail and Voleti, Vikram and Donn{\'e}, Simon and
            Reddy, Mallikarjun Byrasandra Ramalinga and Vainer, Shimon and Boss, Mark},
  year   = {2026}
}