Open-source project
markovka17/dla avatar
markovka17/dla

DLA: the HSE deep learning for audio course, as a public repository

Deep learning for audio processing

769 stars124 forksJupyter NotebookMIT

At a glance

What is it?
DLA is the lecture, seminar and homework material for a graduate deep learning for audio course taught at HSE. It is MIT-licensed and free to work through, but it is course material, not a library you install.
Who is it for?
Work through DLA if you can already write PyTorch and want a graduate-level, self-paced path across speech recognition, separation, TTS and voice biometry, and you are comfortable implementing models from assignments rather than calling a toolkit. Skip it if you need a runnable system now, or if Russian-language lecture videos are a barrier for you.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 9, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What DLA is, and who should work through it

DLA stands for Deep Learning for Audio, and the repository is the teaching material for a course run at the CS Faculty of HSE, with the current version conducted in autumn 2025. It is not a package or a model. It is thirteen weeks of lecture notes, seminar materials and homework assignments, organized into week folders, plus recorded lectures. The audience is anyone who can already write PyTorch and wants a structured path through modern speech and audio deep learning: graduate students, self-learners, or instructors looking for a syllabus to adapt. If you want a ready-to-run toolkit, this is the wrong shelf. If you want to learn to build the models yourself, it is a full curriculum you can follow at your own pace.

Thirteen weeks from signal processing to voice biometry

The syllabus is specific and current. It opens with clean deep learning pipelines, experiment tracking and Hydra in week01, then digital signal processing, Fourier transforms and MFCCs in week02. Automatic speech recognition spans weeks three through five, covering CTC, beam search, language models, LAS and RNN-T. Source separation gets two weeks, naming architectures such as ConvTasNet, DPRNN and the Demucs family. Later weeks move through audio-visual learning, text-to-speech with Tacotron, FastSpeech and HiFi-GAN, neural audio codecs, diffusion-based TTS, and voice biometry with anti-spoofing models including RawNet2 and AASIST. The final week covers explainable AI for deepfake detection. This is a graduate-level sweep of the field, not an introduction to one narrow topic.

Working through the material and the homeworks

There are no install commands here, because there is nothing to install. You clone the repository, open the week folders in order, and each week's own README points at its materials and instructions. The graded work lives in three directories the tree makes visible: hw1_asr for training a speech recognition model, hw3_nv for implementing a neural vocoder TTS model, and project_avss for an audio-visual speech separation project. The homeworks are built on a separate PyTorch project template the README links to, so you set that template up rather than a package from this repository. Lecture recordings are on YouTube, mostly in Russian, with some weeks carrying English recordings in their subdirectories. That language split is a real factor in deciding whether the video material is usable for you.

Where a course repository shows its limits

Treating a course as a product exposes gaps that are not flaws so much as consequences. There is no packaged release, no versioned library, and no API; the default branch is literally named 2025, and past years live as separate branches from 2020 to 2024. The homeworks assume you can stand up the external project template and figure out the environment, because the in-person course supplies teaching assistants and deadlines that a clone does not. The lecture videos being largely in Russian limits the self-study path for those who do not read or listen in that language, even though the written materials are in English. And the material tracks a syllabus, so it teaches methods and expects you to implement them, rather than handing you working reference implementations for every topic.

DLA versus a ready-made speech toolkit

The alternative most people reach for is a speech toolkit such as SpeechBrain or ESPnet. Those are libraries with runnable recipes: you configure a recipe, point it at a dataset, and train a working ASR or TTS system without writing the model yourself. DLA takes the opposite approach on purpose. Its homeworks ask you to implement the pieces, so you come out understanding CTC or a neural vocoder rather than having called a recipe that produced one. If your goal is a working system this week, a toolkit gets you there faster. If your goal is to understand the field well enough to modify or debug those systems, the course path is the one that builds that, at the cost of much more of your time.

MIT license and how current the material is

The repository is under the MIT License, so you can reuse and adapt the materials, including for your own teaching, with attribution; that makes it genuinely usable as a syllabus to fork. On currency, the last push was on 2026-09-16, and the active branch is 2025, so the material reflects a recent run of the course rather than an abandoned archive. The course is maintained by a rotating group of HSE staff listed in the README, and issues are the stated channel for bugs or contribution ideas. Because it is tied to an academic calendar, expect the content to move in yearly jumps as each cohort's branch is prepared, rather than continuous small updates.

Editorial conclusion

Work through DLA if you can already write PyTorch and want a graduate-level, self-paced path across speech recognition, separation, TTS and voice biometry, and you are comfortable implementing models from assignments rather than calling a toolkit. Skip it if you need a runnable system now, or if Russian-language lecture videos are a barrier for you. Start by cloning the repository, reading week01, and setting up the linked PyTorch project template before attempting hw1_asr.

Frequently asked questions

Is DLA a library I can install?

No. DLA is course material: thirteen weeks of lectures, seminars and homework assignments in week folders, not a package or model. You clone the repository and work through it rather than importing it.

What topics does the DLA course cover?

It spans signal processing, automatic speech recognition (CTC, LAS, RNN-T), source separation, audio-visual learning, text-to-speech, neural audio codecs, diffusion TTS, and voice biometry with anti-spoofing and explainable AI.

Are the DLA lectures in English?

The written materials are in English. The recorded lectures on YouTube are mostly in Russian, though some weeks include English recordings in their subdirectories.

Official sources

  1. Issues
  2. License: MIT
  3. markovka17/dla on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/markovka17-dla.svg)](https://hysenlabs.com/projects/markovka17-dla)