← 开源
wyhsirius

LIA

[ICLR 22, TPAMI 24] LIA: Latent Image Animator

Model DevelopmentArchitecturePython
在 GitHub 打开
增长势头
+024 小时新增 Star0.0%
650
Star
68
Fork
+1
本周
2
贡献者
创建于 2022-03-13 · 更新于 2026-10-07 · 今日第 15352 名
主要开发者
README

LIA: Latent Image Animator

Yaohui Wang, Di Yang, François Brémond, Antitza Dantcheva

Project Page | Paper

This is the official PyTorch implementation of the ICLR 2022 paper "Latent Image Animator: Learning to Animate Images via Latent Space Navigation" and TPAMI 2024 paper "LIA: Latent Image Animator".

Replicate

Abstract: Due to the remarkable progress of deep generative models, animating images has become increasingly efficient, whereas associated results have become increasingly realistic. Current animation-approaches commonly exploit structure representation extracted from driving videos. Such structure representation is instrumental in transferring motion from driving videos to still images. However, such approaches fail in case the source image and driving video encompass large appearance variation. Moreover, the extraction of structure information requires additional modules that endow the animation-model with increased complexity. Deviating from such models, we here introduce the Latent Image Animator (LIA), a self-supervised autoencoder that evades need for structure representation. LIA is streamlined to animate images by linear navigation in the latent space. Specifically, motion in generated video is constructed by linear displacement of codes in the latent space. Towards this, we learn a set of orthogonal motion directions simultaneously, and use their linear combination, in order to represent any displacement in the latent space. Extensive quantitative and qualitative analysis suggests that our model systematically and significantly outperforms state-of-art methods on VoxCeleb, Taichi and TED-talk datasets w.r.t. generated quality.

Requirements

  • Python 3.7
  • PyTorch 1.5+
  • tensorboard
  • moviepy
  • av
  • tqdm
  • lpips

1. Animation demo

Download pre-trained checkpoints from here and put models under ./checkpoints. We have provided several demo source images and driving videos in ./data. To obtain demos, you could run following commands, generated results will be saved under ./res.

python run_demo.py --model vox --source_path ./data/vox/macron.png --driving_path ./data/vox/driving1.mp4 # using vox model
python run_demo.py --model taichi --source_path ./data/taichi/subject1.png --driving_path ./data/taichi/driving1.mp4 # using taichi model
python run_demo.py --model ted --source_path ./data/ted/subject1.png --driving_path ./data/ted/driving1.mp4 # using ted model

If you would like to use your own image and video, indicate (source image), (driving video), `` and run

python run_demo.py --model  --source_path  --driving_path 

2. Evaluation

To obtain reconstruction and LPIPS results, put checkpoints under ./checkpoints and run

python evaluation.py --dataset  --save_path 

Generated videos will be save under ``. For other evaluation metrics, we use the code from here.

3. Linear manipulation

To obtain linear manipulation results of a single image, run

python linear_manipulation.py --model  --img_path  --save_folder 

By default, results will be saved under ./res_manipulation.

Acknowledgement

Part of the code is adapted from FOMM and MRAA. We thank authors for their contribution to the community.

BibTex

@inproceedings{
wang2022latent,
title={Latent Image Animator: Learning to Animate Images via Latent Space Navigation},
author={Yaohui Wang and Di Yang and Francois Bremond and Antitza Dantcheva},
booktitle={International Conference on Learning Representations},
year={2022}
}

@ARTICLE{10645735,
  author={Wang, Yaohui and Yang, Di and Bremond, Francois and Dantcheva, Antitza},
  journal={IEEE Transactions on Pattern Analysis and Machine Intelligence}, 
  title={LIA: Latent Image Animator}, 
  year={2024},
  pages={1-16},
}