← 开发者
#4244 4

Alexandre Défossez

@adefossez · France
在 GitHub 打开
5.92k
加权贡献
645
贡献次数
8
项目数
主要项目
上榜的 AI 项目6
1869116249
facebookresearch/audiocraft
+-3

Audiocraft is a library for audio processing and generation with deep learning. It features the state-of-the-art EnCodec audio compressor / tokenizer, along with MusicGen, a simple and controllable music generation LM with textual and melodic conditioning.

23.7k· Jupyter Notebook
33071617
kyutai-labs/moshi
+1

Moshi is a speech-text foundation model and full-duplex spoken dialogue framework. It uses Mimi, a state-of-the-art streaming neural audio codec.

11.2k· Python
5575311
facebookresearch/demucs
+0

Code for the paper Hybrid Spectrogram and Waveform Source Separation

10.4k· Python
6206465
facebookresearch/encodec
+0

State-of-the-art deep learning based audio codec supporting both mono 24 kHz audio and stereo 48 kHz audio.

4.06k· Python
7718779
facebookresearch/denoiser
+0

Real Time Speech Enhancement in the Waveform Domain (Interspeech 2020)We provide a PyTorch implementation of the paper Real Time Speech Enhancement in the Waveform Domain. In which, we present a causal speech enhancement model working on the raw waveform that runs in real-time on a laptop CPU. The proposed model is based on an encoder-decoder architecture with skip-connections. It is optimized on both time and frequency domains, using multiple loss functions. Empirical evidence shows that it is capable of removing various kinds of background noise including stationary and non-stationary noises, as well as room reverb. Additionally, we suggest a set of data augmentation techniques applied directly on the raw waveform which further improve model performance and its generalization abilities.

1.90k· Python
8461953
kyutai-labs/hibiki
+0

Hibiki is a model for streaming speech translation (also known as simultaneous translation). Unlike offline translation—where one waits for the end of the source utterance to start translating--- Hibiki adapts its flow to accumulate just enough context to produce a correct translation in real-time, chunk by chunk.

1.52k· Rust