speech-trident indexes speech models without a licence column
Awesome speech/audio LLMs, representation learning, and codec models
At a glance
- What is it?
- A research group's reading list of speech representation models, neural codecs and speech language models, kept as a single Markdown page in a repository holding nothing else. It carries dates, titles and links, twenty of them to papers and exactly one to code, and no licence anywhere.
- Who is it for?
- Read it as a bibliography rather than a catalogue. It will tell you which speech language models existed in 2025 and roughly what each one claimed, and it will save you the time of assembling that list yourself, which for a field this fragmented is a real saving.
- Can I use it commercially?
- Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
- Is it still maintained?
- Yes. The repository last received commits 86 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 4, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One entry in the whole table links to code
The language model table has four columns: Date, Model Name, Paper Title, Link. In the twenty complete rows dated from 2025-01 to 2025-07, exactly one entry, Fun-ASR-Nano, carries a second link labelled code, pointing at a GitHub repository alongside its paper. Every other row leads to a document and nowhere else. MiniCPM-o is the near exception: its single link does go to a repository, but it is labelled GitHub rather than code. So a page describing itself as a survey of speech language models is, for anyone trying to install one, a bibliography. That is a defensible choice for a research index, and it is also the whole distance between knowing that a model exists and knowing what it would cost to run. Nothing on the page says that implementations were left out on purpose, so a reader has to infer the scope from a single labelled link out of twenty rows.
One column holding preprints, vendored PDFs and a blog
Fifteen of those twenty links resolve to arXiv, spread across three URL shapes: ten abstract pages, four direct PDFs and one HTML rendering. Two more point at PDFs stored inside somebody else's repository, in assets directories, on a default branch in one case and on a branch named cn-readme in the other. MiniCPM-o goes to a repository root. MinMo is labelled Paper and resolves to a project page on a personal site. CSM is labelled blog and resolves to a vendor research post. One column is therefore carrying preprint identifiers, second-hand copies of papers, source trees, project homepages and marketing, and the link labels are the only thing distinguishing them, with paper and Paper both in use. The two vendored PDFs deserve the suspicion. A paper served as a file inside a git repository has none of the durability of a preprint identifier, because the URL names a branch, and a branch can be renamed or removed without the paper moving anywhere.
The model name and the paper title disagree in one row
The Phi-4-Multimodal row carries the title Phi-4-Mini Technical Report, so the name in the second column and the document in the third column describe two differently named models. A reader who copies the first column into a search box asks about one model and is handed a report about another, and then has to work out which way round the mismatch runs before trusting either. The same table puts a plain dash in the Model Name column for the 2025 survey, so a survey paper occupies a row among models with the model column filled by an absence. The CSM row fills the Paper Title column with the phrase Conversational speech generation, which is a description of what the system does rather than the name of a document. Three rows, three ways the columns stop lining up, inside a table that is otherwise strictly formatted.
Two three-part taxonomies that do not overlap
The opening paragraph sorts the field into three areas by layer. Representation models learn structural speech representations which are then quantized into discrete tokens, usually called semantic tokens. Neural codecs learn speech and audio tokens, usually called acoustic tokens, while keeping reconstruction quality and bitrate low. Language models are then trained on top of both. The survey the page points to for what it calls more detailed and technical discussion sorts the same field into three areas by a different axis: pure speech models, speech-aware text models, and models taking speech plus text. One scheme measures where a model sits in a stack, the other measures which modalities it consumes, and neither is a refinement of the other. A reader arriving through the survey and a reader arriving through this repository finish with two incompatible maps of the same field, and the page never says which one to enter by.
A music generator sits inside the speech table
The project description covers speech and audio language models, representation learning and codec models, so the stated scope is wider than speech on purpose. The taxonomy in the opening paragraph is not. The third area is defined as models trained on top of speech and acoustic tokens that show skill in speech understanding and speech generation, and the first thing filed under it is YuE, a foundation model for long-form music generation. Four more rows are multimodal rather than speech-first: Phi-4-Mini assembled from mixture-of-LoRAs, VITA-1.5 aimed at real-time vision and speech, MiniCPM-o covering vision, speech and live streaming on a phone, and MinMo for voice interaction inside a multimodal model. That is five of the twenty rows the opening definition does not really cover, which reads less as an error than as a description that has drifted away from the table it introduces.
No licence column, no licence file, two entries at the root
The repository root holds a README and an assets directory. There is no licence file, and the project's licence field is empty. The table has four columns and none of them is a licence, so a reader scanning twenty speech language models cannot tell which are permissive, which are research-only, and which ship under terms that would stop a product team from shipping them. For a field where the same capability arrives under very different terms, that is the column most worth having. The absence also shapes who the page points at. The contributors panel names six people, each linked to a different personal page on a different host, including a citation profile and a university lab page rather than a project or institution site. The page identifies individuals rather than an organisation that could be asked about reuse, which is consistent with a document maintained by the people who wrote the papers.
One citation under the 2026 heading, and a branch last pushed in July
Two news sections sit above the tables. The 2025 one carries the ten-author landscape survey and a paragraph on what it covers. The 2026 one carries a single BibTeX record for a time-controllable training paper with five authors, plus a pointer to a separate repository surveying spoken dialogue models. Two names appear in both citations, and both are on the contributors panel, so the newer item comes from the same group that wrote the older survey and the direction of travel is towards one survey repository per subtopic. Maintenance is quiet rather than abandoned. The project has no releases at all, so there is no version history to diff a page against, and the last push to the default branch is dated 2026-07-10. The newest row in the language model table carries the date 2025-07, so a reader landing on the page now has no way to tell which rows were checked recently and which were filed in a hurry.
Editorial conclusion
Read it as a bibliography rather than a catalogue. It will tell you which speech language models existed in 2025 and roughly what each one claimed, and it will save you the time of assembling that list yourself, which for a field this fragmented is a real saving. It will not tell you which of them you are allowed to ship, which is the question that decides most adoptions, and it will not tell you which are still maintained, because the page has no licence and no release history to answer either. Before building on anything here, open the model you actually want, read its own licence and its own repository, and treat every row as a pointer to a claim rather than a description of software.
Frequently asked questions
What does the speech-trident repository survey?
Three areas: speech representation models that quantize speech into semantic tokens, neural codec models that learn acoustic tokens at low bitrate while keeping reconstruction quality, and speech language models trained on top of both tokens for speech understanding and generation.
What columns does the speech-trident model table have?
Four: Date, Model Name, Paper Title and Link. There is no licence column and no code column, so the table cannot answer what a listed model costs to run or whether it can be shipped.
Does speech-trident link to code for any model it lists?
One. Fun-ASR-Nano carries a second link labelled code to a GitHub repository alongside its paper. MiniCPM-o links to a repository as well, but under the label GitHub rather than code.
How does speech-trident differ from the speech language model survey it links?
The repository sorts models by layer, separating representation, codec and language models, while the survey sorts them by modality composition, separating pure speech, speech-aware text, and speech plus text models. The page offers both without mapping one onto the other.
How is speech-trident maintained?
Quietly. It has no GitHub releases, so there is no version history, and the last push to the default branch is dated 2026-07-10. It has 1248 stars, 74 forks and 6 open issues, and the repository root holds only the README and an assets directory.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ga642381-speech-trident)