MASCOT reports four headline numbers with no baseline, and its install stops mid-string
EMNLP 2026 Main Conference Paper: MASCOT: Towards Multi-Agent Socio-Collaborative Companion Systems (https://arxiv.org/abs/2601.14230)
At a glance
- What is it?
- The EMNLP 2026 companion-agent framework from hello-diana, two GRPO stages over Qwen3 models with LoRA adapters. The pipeline details are unusually precise. The measurements, the install command and one licence reference are not.
- Who is it for?
- This is a research artefact rather than a product, and it is honest about that in the places that matter most. It tells you which model plays which role, how many candidates each query gets, how many preference pairs that yields before filtering, and it explicitly marks the offline DPO path as not the paper reproduction path, which is the kind of warning most repositories omit.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 42 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The results table mixes three kinds of number and names no baseline
Four figures appear under Results, and they are not four measurements of the same thing. Persona Consistency is up 14.1 and Social Contribution is up 10.6, with no units and nothing named as the thing they are up from. Human preference head to head is 69%, with no number of raters, no number of comparisons and no statement of what the alternative system was. The last row is not a result at all: trainable policy parameters at 0.187% with LoRA at rank 16, which is a property of the training setup rather than an outcome. Below, the evaluation is described as human, multi-judge with GPT-4o, Gemma-3-27B and Phi-4, and automatic metrics across Empathetic Dialogues and QMSum, plus out-of-domain ESConv. No dataset size, baseline system, seed count or spread appears anywhere on the page.
The installation section ends in the middle of a command
Installation is the first thing a reader tries, and the visible text stops inside it:
pip install -e . # core: episode generation + rubric evaluation
pip install -e ".The first line installs the core, which the comment defines as episode generation plus rubric evaluation. The second line opens a quoted extras specification and stops. What should follow is recoverable from the packaging metadata rather than from the page, which defines five optional groups: train with torch, trl, peft and accelerate for reward model, GRPO and director work, serve with vllm for the local model server, metrics with torch and nltk for the automatic non-LLM scores, wandb, and api-judges for the Anthropic and Google judges. So the page's install section cannot be completed as written.
The packaging metadata points at a licence file the repository does not contain
The project metadata carries a licence comment that is worth reading carefully. Apache-2.0 applies to the source code, and the comment says to see THIRD_PARTY_NOTICES.md for the licence scope of figures, website assets, datasets and third-party components. The repository root does not have that file. What it has is a licence file, a readme, the configuration directory, the media directory, an index.html, the packaging file, a scripts directory and the source tree. So the one pointer that would tell you what licence the figures, the media assets and the evaluation datasets carry resolves to nothing in this tree. The datasets named on the page are third-party corpora, and the figures and site assets belong to the project page, so the scope question is real and the file meant to answer it is absent.
Eight candidates become twenty-eight pairs, with two reward models of one size
Stage 1 is spelled out with numbers. K equals 8 candidate responses per context and persona, an LLM judge scores them with rubric feedback, and preference pairs are built from the scored candidates, up to K times K minus one over two, which is 28 unordered pairs per query before margin filtering. A scalar persona reward model of Qwen3-0.6B is trained on those pairs, and the Speaker policy, Qwen3-8B, is optimized with GRPO against it plus small format and conciseness rewards. Stage 2 repeats the shape: a group reward model of Qwen3-0.6B scores full trajectories and the Director, also Qwen3-8B, is optimized with GRPO over them. Two small judges and two large policies, with LoRA adapters only and base weights referenced by Hugging Face name.
One section is explicit about which command does not reproduce the paper
The Director has two training modes and the page distinguishes them carefully, which is unusual and worth crediting. `train-director-grpo` performs online, group-relative optimization over trajectory rewards, is the method described in the paper, and is the default. `train-director-dpo` performs offline preference optimization over stored directive and trajectory pairs produced by the branching episode generator, is kept as an optional alternative, and is stated in italics not to be the paper reproduction path. The same care appears in the highlights, which note that reward-model checkpoints and explicitly merged checkpoints may contain full model weights, unlike the LoRA-only adapters that reference base weights by name. A reader knows which command to run and which artefacts are heavier than the headline suggests.
Personas live in two files and the rosters differ by domain
The agent rosters are fixed sets, not generated ones. Emotional support runs three personas: an emotional validator called The Anchor, an action-oriented guide called The Catalyst, and a growth advocate called The Beacon, evaluated on Empathetic Dialogues and out-of-domain ESConv. Meetings runs four: Minutes Scribe, Decision Logger, Action Item Captain and Critic, evaluated on QMSum. So the candidate count of 8 applies per persona, and the two domains do not use the same number of speakers. The definitions are not in one place: personas and prompts live in a configuration module under `config/` and in a templates module under `src/prompts/`. Simulated users optionally carry Big Five personality profiles from a third module, which is how the framework varies the human side of a conversation.
A training pipeline that needs a model server, not a companion application
What the repository delivers is described as a self-contained pipeline: episode generation, rubric judging, preference data construction, reward modeling, GRPO and DPO training, and evaluation. All of it runs against an OpenAI-compatible vLLM server that is served separately, which is why the serve extra exists and why the core dependency comment says the server is served elsewhere. Core dependencies are openai, datasets, transformers, pandas, numpy, pydantic, python-dotenv, pillow and tqdm, which is enough to generate episodes and run rubric evaluation and nothing more. There is no chat interface, no packaged application and no release: the project version in the metadata is 0.1.0, the repository publishes no releases, and the last commit is dated 24 August 2026.
Editorial conclusion
This is a research artefact rather than a product, and it is honest about that in the places that matter most. It tells you which model plays which role, how many candidates each query gets, how many preference pairs that yields before filtering, and it explicitly marks the offline DPO path as not the paper reproduction path, which is the kind of warning most repositories omit. What a reader cannot do is reproduce or evaluate the claims from the page. The results are deltas and a preference win rate with no baseline, no sample size and no variance, the install command stops mid-string, and the packaging metadata defers dataset and figure licensing to a notices file that the repository root does not contain. Before building on it, read the paper for the baseline and the evaluation protocol, install from the packaging metadata rather than the page, and settle the licence question for Empathetic Dialogues, QMSum and ESConv yourself, since the pointer you are given does not resolve. It suits a researcher reproducing a two-stage training recipe. It does not suit someone expecting a working companion application.
Frequently asked questions
What is MASCOT?
A multi-agent framework for building socio-collaborative companions, published as an EMNLP 2026 main conference paper with a project page and an arXiv entry. It has two stages: persona-aware behavioural alignment of individual speaker agents, then collaborative dialogue optimisation where a director agent chooses who speaks next and issues natural-language directives. It targets interaction quality in affective and collaborative settings rather than task efficiency.
Which models does MASCOT train and judge with?
Qwen3-0.6B serves as both the persona reward model and the group reward model, and Qwen3-8B as both the Speaker policy and the Director policy. Only LoRA adapters are trained, with base weights referenced by Hugging Face name, and the reported share of trainable parameters is 0.187% at rank 16.
How do I install MASCOT?
The core install is `pip install -e .`, described as enough for episode generation and rubric evaluation. The second documented install line stops inside a quoted extras specification, so the extras have to be read from the packaging metadata instead, which defines train, serve, metrics, wandb and api-judges groups. Python 3.10 or newer is required.
Does MASCOT ship a companion chat app?
No. The repository is a training and evaluation pipeline: episode generation, rubric judging, preference data construction, reward modeling, GRPO and DPO training, and evaluation, all against an OpenAI-compatible vLLM server that has to be served separately. There is no chat interface, and the repository publishes no releases, with the project version at 0.1.0.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/hello-diana-mascot)