읽고 쓴 글.
논문 리뷰와 공부하면서 적은 글을 모았습니다. 제목이나 주제로 찾아볼 수 있습니다.
RawBoost: A Raw Data Boosting and Augmentation Method applied to Automatic Speaker Verification Anti-Spoofing
RawBoost
Aasist: Audio Anti-Spoofing Using Integrated Spectro-Temporal Graph Attention Networks
AASIST
VALL-E 2: Neural Codec Language Models are Human Parity Zero-Shot Text to Speech Synthesizers
VALL-E 2
VALL-E R: Robust and Efficient Zero-Shot Text-to-Speech Synthesis via Monotonic Alignment
VALL-E R
Speak Foreign Languages with Your Own Voice: Cross-Lingual Neural Codec Language Modeling
VALL-E X
Natural language guidance of high-fidelity text-to-speech with synthetic annotations
Parler-TTS
Should you use a probabilistic duration model in TTS? Probably! Especially for spontaneous speech
TTS에서 probabilistic duration model이 효과적인 상황 탐구
Imaginary Voice: Face-styled Diffusion Model for Text-to-Speech
사람의 얼굴에 어울리는 목소리를 자동으로 만들어주는 디퓨전 기반 TTS 모델
NN-KOG2P: A Novel Grapheme-To-Phoneme Model for Korean Language
한국어의 발음특성을 고려한 FFNN G2P 모델
Grad-TTS: A Diffusion Probabilistic Model for Text-to-Speech
Diffusion Probabilistic Model 기반 decoder를 사용한 TTS 모델
Label Embedding for Chinese Grapheme-to-Phoneme Conversion
contrastive learning을 활용한 중국어 G2P
Almost Unsupervised Text to Speech and Automatic Speech Recognition
Transformer로 TTS와 STT를 동시에 하는 방법
Neural Lexicon Reader: Reduce Pronunciation Errors in End-to-end TTS by Leveraging External Textual Knowledge
G2P 없이 Lexicon 만으로 TTS 발음 오류를 줄여보자
Multilingual grapheme-to-phoneme conversion with byte representation
byte로 다중 언어 G2P 성능을 높여보자
Unified Mandarin TTS Front-end Based on Distilled BERT Model
BERT에서 뽑아낸 feature로 중국어 TTS 전처리 과정 단순하게 만들기
A bi-directional LSTM approach for polyphone disambiguation in mandarin chinese
중국어 G2P를 위한 bi-directional LSTM
Sequence-to-sequence neural net models for grapheme-to-phoneme conversion
Seq2Seq 기반 G2P
Grapheme-to-phoneme conversion using Long Short-Term Memory recurrent neural networks
G2P에 LSTM을 처음 적용한 논문
Style Tokens: Unsupervised style modeling, control and transfer in end-to-end speech synthesis
style embedding을 통해 Tacotron으로 합성한 음성의 스타일을 조절해보자
LDNet: Unified Listener Dependent Modeling in MOS Prediction for Synthetic Speech
MBNet에서 단점을 파악하고 개선해보자
MBNet: MOS Prediction for Synthesized Speech with Mean-Bias Network
평가자 정보를 활용해 더 정확하게 MOS를 예측해보자
FastPitchFormant: Source-Filter Based Decomposed Modeling for Speech Synthesis
FastPitch에 source-filter 이론을 접목시켰다
FastPitch: Parallel text-to-speech with pitch prediction
FastSpeech2와 거의 동시에 나온 FastSpeech 후속 버전
Meta-StyleSpeech: Multi-Speaker Adaptive Text-to-Speech Generation
TTS에 meta learning 적용하기
HiFi-GAN: Generative Adversarial Networks for Efficient and High Fidelity Speech Synthesis
GAN 기반 보코더의 한계를 극복
Parallel Tacotron 2: A Non-Autoregressive Neural TTS Model with Differentiable Duration Modeling
Parallel Tacotron + duration model
Parallel Tacotron: Non-autoregressive and controllable TTS
Non Autoregressive Tacotron + VAE
AdaSpeech 3: Adaptive text to speech for spontaneous style
AdaSpeech를 자연스러운 발화가 가능하도록 만들어보자
AdaSpeech 2: Adaptive text to speech with untranscribed data
AdaSpeech에서 전사가 안된 데이터를 활용해서 TTS를 할 수 있게 만들어보자
Wave-Tacotron: Spectrogram-free end-to-end text-to-speech synthesis
스펙트로그램 없이 end-to-end TTS 하기
Glow-TTS: A generative flow for text-to-speech via monotonic alignment search
Flow 기반의 TTS 모델
Melgan: Generative adversarial networks for conditional waveform synthesis
GAN 기반의 빠르고 효율적인 보코더
Parallel WaveGAN: A fast waveform generation model based on generative adversarial networks with multi-resolution spectrogram
GAN loss와 새로운 loss를 결합하여 TTS를 하는 모델
FastSpeech 2: Fast and high-quality end-to-end text to speech
FastSpeech에서 teacher forcing을 제거한 모델
FastSpeech: Fast, robust and controllable text to speech
Transformer 기반의 빠르고 조절 가능한 TTS 모델
Informer: Beyond Efficient Transformer for Long Sequence Time-Series
긴 시계열 예측에 특화된 트랜스포머 모델
Shape and Time Distortion Loss for Training Deep Time Series Forecasting Models
비정상성을 가진 시계열 데이터를 위한 새로운 Loss function 제안
N-BEATS: Neural Basis Expansion Analysis for Interpretable Time Series Forecasting
오직 딥러닝 아키텍처만을 이용한 TS모델
Temporal Fusion Transformers for Interpretable Multi-horizon Time Series Forecasting
트랜스포머를 활용한 multi-horizon time-series 예측
Deep Representation Learning in Speech Processing: Challenges, Recent Advances, and Future Trends
음성처리 서베이 논문
아직 일치하는 글이 없어요.
다른 검색어나 주제를 선택해 보세요.