MELD, the multi-party multimodal emotion dataset built on Friends
MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation
At a glance
- What is it?
- MELD extends the EmotionLines dataset with audio and visual streams for more than 1,400 multi-speaker dialogues drawn from Friends, and it is still the reference benchmark for emotion recognition in conversation. Its severe class imbalance and its single-show provenance are the two facts that should shape how you use it.
- Who is it for?
- Use MELD when you need a standard benchmark for emotion recognition in conversation with several speakers and all three modalities, and report macro-averaged scores rather than accuracy, because Neutral outnumbers Disgust by more than an order of magnitude in the training split. Do not use it to claim general conversational emotion understanding, since every utterance comes from a single English-language sitcom.
- Can I use it commercially?
- Yes, with conditions. GPL-3.0 is a copyleft licence: if you distribute software that includes it, you must release that software's source code under the same licence. Running it internally without distributing it does not trigger that obligation.
- Is it still maintained?
- Yes. The repository last received commits 142 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What MELD adds on top of the EmotionLines dialogues
EmotionLines gave the field a large set of dialogues with emotion labels, and it is text only. That is its limitation, and it is the gap MELD was built to close. MELD keeps the same dialogue instances but adds the audio and visual streams alongside the text, giving every utterance three aligned modalities.
The scale is modest and that is deliberate. MELD contains more than 1,400 dialogues and about 13,000 utterances, all drawn from the Friends television series, with multiple speakers per dialogue rather than a fixed pair. Each utterance carries one of seven emotion labels: Anger, Disgust, Sadness, Joy, Neutral, Surprise and Fear. On top of that, every utterance has a sentiment label of positive, negative or neutral.
The research case the README makes is about context, not about modalities for their own sake. Emotion shifts as a conversation runs, so modelling context across turns is the hard part, and the hypothesis is that having audio and visual evidence for each turn improves that modelling. It also positions the dataset against IEMOCAP and SEMAINE, which the README describes as multimodal and labelled but dyadic, meaning limited to two participants and therefore not a test of multi-party understanding. For a working researcher, the practical reason to care is that MELD became the standard reported benchmark, which is what makes your numbers comparable to published work at all.
The split table, and what 4,003 emotion shifts imply for context modelling
The statistics table in the README is the part of this dataset most worth reading before you write a loader, because the splits are uneven in ways that affect how you evaluate.
The dialogue split is 1,039 train, 114 dev and 280 test, and the utterance split is 9,989, 1,109 and 2,610. Unique word counts are 10,643, 2,384 and 4,361, so the vocabulary in the test split is not simply a subset of the training vocabulary at the same rate as its size. Average utterance length sits at about eight words, with a maximum of 69 in training and lower maxima in dev and test, and average utterance duration is 3.59 seconds in training, 3.59 in dev and 3.58 in test, which is unusually consistent and suggests the splits were not made along a time boundary.
The number that explains the difficulty is emotion shift: 4,003 in training, 427 in dev and 1,003 in test. That is roughly 40 percent of training utterances sitting at a point where the speaker's emotion changes from the previous turn. A model that reads utterances in isolation is therefore wrong by design on a large slice of the data, and a dialogue-context model has to propagate information across turns without over-smoothing a speaker who genuinely changes mood. Speaker counts of 260, 47 and 100 give you another axis: how many distinct voices your model has to generalise across within a split.
Neutral against Disgust is the statistic that decides your metric
The dataset distribution table is where anyone training on MELD should start, because the classes are not remotely balanced. In the training split, Neutral accounts for 4,710 utterances. Disgust accounts for 271, Fear 268, Sadness 683, Anger 1,109, Surprise 1,205 and Joy 1,743. The ratio between the largest and smallest class is better than seventeen to one, and the same shape holds in dev and test.
That has a direct consequence. A model that predicts Neutral for everything scores about 47 percent accuracy on the training distribution and lands in the same region on the test split, because Neutral is 1,256 of 2,610 there. Accuracy on MELD is therefore close to meaningless as a headline number, and any paper reporting it without a majority-class baseline beside it is telling you very little. The interesting question is what happens on Disgust and Fear, which together are under six percent of the training data.
A second consequence is about iteration. With 268 training examples for Fear, a large model's extra capacity has little to fit, and aggressive augmentation of the rare classes will change the distribution you are pretending to measure. The README does not document a recommended metric or a standard split protocol beyond the counts, so the reporting convention has to come from the papers you compare against, which is another reason the baseline repository the README links to matters more than the code in this tree.
Subtitle timestamps, two constraints, and a dialogue count that no longer matches EmotionLines
The dataset creation section explains how alignment was achieved, and it explains an inconsistency you will hit if you try to join MELD back to its parent dataset.
The first step was finding the timestamp of every utterance. The authors crawled the subtitle files of the episodes, which contain the beginning and end timestamp of each utterance, which yielded a season ID, an episode ID and a timestamp for every turn. Two constraints were imposed while collecting those timestamps: the timestamps within a dialogue must be in increasing order, and every utterance in a dialogue must belong to the same episode and scene.
Those constraints turned out to be load-bearing. Applying them revealed that a few EmotionLines dialogues actually consist of several natural dialogues concatenated together, and those cases were filtered out of MELD. The README is explicit about the consequence: because of that error correction step, MELD has a different number of dialogues compared to EmotionLines. So if you are reproducing a paper that reports results on both, the two dialogue sets are not the same set, and any per-dialogue join keyed on EmotionLines identifiers will not line up. Treat MELD as its own dataset with its own counts, not as EmotionLines plus extra columns.
Having the timestamps, the authors then extracted the corresponding audio and visual clips for each utterance. That is where the multimodal part comes from, and it is also why the dataset is distributed as media rather than as a table of pre-extracted features.
Where the data lives, and the three sibling repositories you will also need
The README is unambiguous that the data is not in this repository. For downloading it, the README points to a dataset page on Hugging Face:
https://huggingface.co/datasets/declare-lab/MELDThere are no install steps in the README, no loader script to copy and no pip package to add. What the repository itself holds is a top level of `LICENSE`, `README.md`, and the `baseline/`, `data/`, `utils/` and `images/` directories, none of whose contents the README describes in the excerpt available here. The project page at affective-meld.github.io is given as the place for more details.
Three related repositories matter as much as this one for anyone reproducing results. The first is where the visual features were released: the README says visual features extracted using Resnet are available through the declare-lab group. If your baseline consumes pre-extracted visual vectors rather than raw video, that is the artefact you want, and it is not in this tree. The second is the baseline repository, which the README points to for updated baselines, and which also contains the COSMIC directory with the code for the 10/10/2020 state of the art on this dataset, described in the README as COmmonSense knowledge for eMotion Identification in Conversations, with the paper on arXiv.
The third relationship is historical. The README also links an unrelated project from the same group on IQ testing of large language models, which is a useful signal about where the lab's attention has moved. MELD itself has not been extended: the newest dated entry in the updates list is 10/10/2020, before that an ACL 2019 acceptance on 22/05/2019 and the release of Dyadic MELD the same day for testing dyadic conversational models.
MELD against IEMOCAP, SEMAINE, and EmotionLines
The alternatives are real and the choice comes down to structure, not quality.
IEMOCAP and SEMAINE are the two datasets the README itself names as comparisons. Both are multimodal and both carry emotion labels per utterance. Both are dyadic, which is the disqualifier for multi-party work: a model validated only on two-person conversation has not been tested on the shift patterns that appear when a third and fourth speaker join, and MELD's speaker counts of 260, 47 and 100 are exactly that test. EmotionLines is the other neighbour, and it remains the right choice for a text-only baseline, since it is the dataset MELD was derived from and carries the same dialogue lineage without the media.
The trade-off runs the other way once you leave the benchmark world. MELD's dialogues all come from one American situation comedy, in one language, with the acting conventions, the pacing and the humour that implies. A model trained on it learns Friends. Whether it transfers to customer support calls, to meetings, or to another language is an open question that this dataset cannot answer, and no leaderboard entry will tell you. The same single-domain problem applies to the speaker axis: 260 distinct training speakers is a good number for generalising across voices, but all of them are performers playing characters.
There is also a leaderboard, and it is a picture. The README embeds the results as an image file, so the numbers are not machine-readable, not sortable, and not linkable to a specific configuration. The dated updates and the pointer to the baseline repository are the more reliable record of what state the art was at any moment.
GPL-3.0, no releases, and a silence that starts in 2020
The licence is GPL version 3, with the text in the `LICENSE` file at the repository root. For a dataset repository that covers the code, the baseline implementations and any files distributed alongside them. What it does not settle is the status of the underlying material, and the README is silent on that point: the clips were extracted from Friends episodes, and the repository says nothing about the rights to the source footage or about what a downstream user may redistribute. Anyone planning to publish a derived corpus should resolve that before publishing, and the answer will not be in this README.
The maintenance record is thin in a way that suits a finished dataset. There are no GitHub releases, so there is no tag to pin and no release notes to read. The last push to the repository was on 2026-05-17, and the repository is not archived, but the README's own update log stops at 10/10/2020 and the earlier entries are from 2018 and 2019, including a fix on 15/11/2018 for a problem in `train.tar.gz` that suggests an earlier distribution channel. The current distribution path is the Hugging Face dataset page, and the README does not document a sync process between the two.
So the honest summary is that the data is stable and the tooling is elsewhere. A dataset that has not needed a format change in six years is a good sign for reproducibility, and a repository whose leaderboard is a PNG and whose best baselines live in a sibling project is a sign to budget your time around the sibling project.
Editorial conclusion
Use MELD when you need a standard benchmark for emotion recognition in conversation with several speakers and all three modalities, and report macro-averaged scores rather than accuracy, because Neutral outnumbers Disgust by more than an order of magnitude in the training split. Do not use it to claim general conversational emotion understanding, since every utterance comes from a single English-language sitcom. Verify first by downloading from the dataset page on Hugging Face rather than from the repository's data directory, checking the split counts against the table in the README, and reading the COSMIC paper the README points to for the baseline you intend to beat.
Frequently asked questions
Where can I download the MELD dataset?
The README directs you to the dataset page on Hugging Face at https://huggingface.co/datasets/declare-lab/MELD. The repository itself is not the download location, and it gives no install or loader script; the project page at affective-meld.github.io is listed for further details.
How large is the MELD dataset?
The README states more than 1,400 dialogues and 13,000 utterances from the Friends series, split into 1,039 training, 114 dev and 280 test dialogues, and 9,989, 1,109 and 2,610 utterances respectively. Average utterance duration is about 3.59 seconds in all three splits.
Which emotion labels does MELD provide?
Each utterance is labelled with one of seven emotions: Anger, Disgust, Sadness, Joy, Neutral, Surprise and Fear. Every utterance also carries a sentiment label of positive, negative or neutral. The classes are heavily imbalanced, with Neutral at 4,710 training utterances against 271 for Disgust and 268 for Fear.
What are the pre-extracted visual features for MELD?
The README notes that visual features extracted using Resnet have been released, and links to the declare-lab group for them. Those features are not part of this repository, so a baseline that consumes visual vectors rather than raw video needs a second download.
Where can I find current baselines for MELD?
The README points to a separate baseline repository in the declare-lab group for updated baselines, and specifically to its COSMIC directory for the code behind the 10/10/2020 result described as COmmonSense knowledge for eMotion Identification in Conversations, with the paper available on arXiv.
What licence is the MELD dataset released under?
The repository is licensed under GPL version 3, with the text in the LICENSE file at the root. The README does not address the licensing of the underlying television material the clips were extracted from, which is a separate question for anyone redistributing derived data.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/declare-lab-meld)