GALA: Gaussian Animation via Linear Approximation

One Basis to Animate Them All

Gaussian Blendshape Distillation for Real-Time Avatars

TL;DR Gaussian avatars rerun a heavy network for every new pose. GALA distills it into a linear blend of shared blendshapes and a shallow MLP: up to ×2,659 faster on CPU, without retraining it.

  • ×2,659the largest speedup of CPU animation
  • 60 fpsin a phone’s web browser
  • 1 basisper model, shared by all identities
  • 0host models retrained

Abstract

3D Gaussian avatars support fast rendering, however, their real-time animation is often challenged by the costly neural inference. We address this bottleneck and show that the animation of pretrained avatar models can be closely approximated by a linear combination of identity-independent blendshapes. Building on this finding, we introduce GALA (Gaussian Animation via Linear Approximation), a distillation method that replaces per-frame heavy neural decoding with a shallow coefficient predictor and a linear blend. To improve fidelity and reduce memory requirements, we propose to construct the basis using block-local PCA under a rendering-aware metric and a memory budget. Our method learns a shallow MLP network to predict blendshape coefficients and applies to various animation architectures without retraining original models. We validate GALA by accelerating the inference of three distinct avatar models for 3D animation of facial expressions and full-bodies with clothing dynamics. Across these models, our distillation generalizes to held-out identities and reduces CPU animation cost by up to three orders of magnitude while preserving most of the rendering quality. Excellent results of our method confirm the shared linear structure of learned avatar representations and enable highly efficient and accurate animation at frame rates reaching up to 60fps on mobile devices.

Gaussian avatars render fast. Animating them doesn’t.

New viewpoint rasterize only

New pose rerun the transformer

A new viewpoint is cheap: 622 fps. A new pose reruns a large network: 1.8 s, even on a desktop GPU.

DynaAvatar on an RTX 6000 Ada.

Animation is a linear combination of basis vectors.

Take the residual between the animated avatar and the neutral avatar of the same identity.

r_i = A(w_i, theta_i) - A(w_i, theta_0)

It is closely approximated by a weighted sum of basis vectors.

r_i is approximately c_i1 u_1 + c_i2 u_2 + ... = U c_i
basis vectors
the same for all identities, even unseen ones
weights
depend on the identity and the pose

It even holds for identities never seen before.

Inputone generated image

FlexAvatar hostoriginal animation

Neutral onlyno blendshapes

Projectedonto the shared basis

Blendshapes are local, and shared.

The same blendshapes close an eye, close a mouth or lift a jacket on every identity. Drag the slider.

− +

First column: where they act. Each row moves a group of local blendshapes between two states.

From a pretrained model to a linear blend.

Overview of GALA: host model, basis construction, coefficient distillation, and inference once per subject and every frame.

Three steps, and the host is never retrained.

Build one basis from the host’s residuals, train a shallow network to predict the coefficients, then blend. Click a stage.

A neural network recomputes the avatar at every frame.

We run the host on many identities and poses, and keep each residual to the neutral avatar.

Many residuals, one basis: local, rendering-aware PCA.

All residuals go into one matrix. A PCA over local blocks, in a rendering-aware metric and under a memory budget, gives the basis U.

Same memory budgetPSNR to host, AGORA
Simple PCA32.8 dB
Ours42.6 dB

A shallow network learns the coefficients.

A shallow MLP predicts the coefficients from the identity and the driving signal.

Every frame is a tiny network and a linear blend.

GALA closely matches each host.

Expressions, mouth interiors and garment motion, at a few milliseconds per frame.

Host+GALA (ours)Error
154 ms / frame5.5 ms / frame0 – 0.2
Host+GALA (ours)Error
234 ms / frame4.4 ms / frame0 – 0.2
Host+GALA (ours)Error
42.9 s / frame16.1 ms / frame0 – 0.2

Hosts: AGORA (Fazylov et al.), FlexAvatar (Kirschstein et al., CVPR 2026), DynaAvatar (Kwon et al., CVPR 2026).

Ahead of the state of the art on most metrics.

GALA(AGORA)

on FFHQ

3 of 5

metrics best, 1 second best

FIDAEDAED-jawIDAPD

vs Next3D, GAIA, EG3D, GGHead

GALA(FlexAvatar)

on VFHQ and Ava256

8 of 12

metrics best, 2 second best

PSNRLPIPSAEDCSIMAEDAPD
PSNRAKDCSIMPSNRAKDCSIM

vs LAM, GAGAvatar, Portrait4D-v2, GPAvatar, Avat3r

GALA(DynaAvatar)

on 4D-Dress

3 of 3

metrics best

PSNRSSIMLPIPS

vs IDOL, LHM, PERSONA

Best2nd bestOther Quality metrics from the paper; the host itself is not ranked.

All three hosts, live in a phone’s web browser.

  • computed on the phone
  • web browser
  • iPhone 15
AGORA + GALA 60 fps
FlexAvatar + GALA 60 fps
DynaAvatar + GALA 60 fps

WebGL only: the network and the blend run as shader passes. No host network runs in the browser.

A whole table of avatars…

  • iPhone 15
  • 5 heads
  • 40–50 fps

…or a crowd.

  • iPhone 15
  • 5 avatars
  • 20 fps

BibTeX

If you use our work, please cite it as:

@article{fazylov2026gala,
  title   = {One Basis to Animate Them All: Gaussian Blendshape Distillation for Real-Time Avatars},
  author  = {Fazylov, Ramazan and Lefkimmiatis, Stamatis and Laptev, Ivan},
  journal = {arXiv preprint},
  year    = {2026}
}