Chenxi
- zai-org/CogVideo
- TIGER-AI-Lab/AnyV2V
- microsoft/Bringing-Old-Photos-Back-to-Life
- showlab/Show-1
- magic-research/piecewise-rectified-flow
- google-research/maxim
- sczhou/Upscale-A-Video
- bytedance/res-adapter
- declare-lab/tango
- deepseek-ai/DeepSeek-Math
[NeurIPS 2022] Towards Robust Blind Face Restoration with Codebook Lookup Transformer
SwinIR: Image Restoration Using Swin Transformer (official repository)
Latent Consistency Models: Synthesizing High-Resolution Images with Few-Step Inference
人像卡通化探索项目 (photo-to-cartoon translation project)
[CVPR 2022] Thin-Plate Spline Motion Model for Image Animation.
DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
The state-of-the-art image restoration model without nonlinear activation functions.
FILM: Frame Interpolation for Large Motion, In ECCV 2022.
LLaMA-Omni is a low-latency and high-quality end-to-end speech interaction model built upon Llama-3.1-8B-Instruct, aiming to achieve speech capabilities at the GPT-4o level.
Image to prompt with BLIP and CLIP
[IJCV2024] Exploiting Diffusion Prior for Real-World Image Super-Resolution
Code for ALBEF: a new vision-language pre-training method