The MoE survey repository is a dated index, not a reading list
[TKDE'25] The official GitHub page for the survey paper "A Survey on Mixture of Experts in Large Language Models".
At a glance
- What is it?
- A-Survey-on-Mixture-of-Experts-in-LLMs holds a README, a licence and an assets folder, and the README is the index: papers in reverse chronological order, each with an arXiv link, a venue tag and a release date, with open-source models above an arrow and proprietary ones below it. The survey itself was accepted by TKDE and lives on arXiv at 2407.06204.
- Who is it for?
- Use this repository as a map of the mixture-of-experts literature rather than as a text to read through. Its value is that it is dated and categorised, so you can see what changed in a given month, which venues are publishing MoE work, and which models are open.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 53 days ago.
- What is it written in?
- GitHub does not report a main language for this repository.
Answers come from the project's GitHub data, last synced on October 10, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three entries at the top level, and the paper is not one of them
The repository is unusually small for what it circulates. At the top level there are three things: a LICENSE, a README.md and an assets folder. There is no source directory, no notebook, no dataset and no build, and GitHub reports no primary language for the repository at all, which is consistent with a page whose content is prose plus images. The work itself is the paper, titled A Survey on Mixture of Experts in Large Language Models, and the repository is described in its own metadata as the official GitHub page for it. The paper sits on arXiv at 2407.06204, and the README points readers there for the detailed contents rather than summarising the argument. The acceptance news is stated as an important notice near the top: the survey has been accepted by TKDE, the IEEE Transactions on Knowledge and Data Engineering journal. For a reader deciding where to spend an afternoon, that distinction matters, since the README is an index and the article is the substance, and the two are maintained on different clocks.
Open-source models sit above the arrow, proprietary ones below it
The organising device is a timeline image, and the caption explains the encoding precisely. The timeline is structured primarily by the release dates of the models, models placed above the arrow are open-source, and those below the arrow are proprietary and closed-source. That single split is the most decision-relevant piece of metadata on the page, because it is not something a bibliographic index usually carries: a reader who needs to build on a result and a reader who only needs to cite it want opposite halves of the same list. Colour carries the second dimension, with four domains marked distinctly. Natural Language Processing is green, Computer Vision is yellow, Multimodal is pink, and Recommender Systems is cyan. The recommendation-systems category is the one worth pausing on, since a survey titled for large language models also tracks MoE work in recommender systems, which tells you the scope is the sparse-expert architecture rather than the language model. The previous version of the page is dated January 2025, so the timeline has been extended at least once since the survey first circulated. The taxonomy itself is presented as a centred figure rather than as prose in the file, which is why the paper list carries the detail a reader can actually search, and why the table of contents lists Taxonomy separately from the paper list rather than nesting it.
Venue tags do the filtering: ICML, CVPR, ICLR, AAAI, ASPLOS, ISCA
Every entry carries a venue tag in square brackets and an explicit date, and reading the tags across the list is more informative than any single paper. The dates are also written out in full as year, month and day rather than as a month label, which makes it visible that two entries sharing a date were posted on different days. The 2026 entries name ICML, CVPR, ICLR, AAAI, ASPLOS and ISCA alongside a large number of arXiv preprints, which says two things at once: the architecture has reached the main artificial intelligence and vision conferences, and it has also reached the systems venues, which is where the deployment papers live. A few examples make the spread concrete. GEM: Graph-Enhanced Mixture-of-Experts with ReAct Agents for Dialogue State Tracking is tagged AAAI 2026. LAER-MoE: Load-Adaptive Expert Re-layout for Efficient Mixture-of-Experts Training is tagged ASPLOS 2026. MoE-Hub, on hardware-accelerated communication for MoE overlap on multi-GPU systems, and Accelerating MoE with Dynamic In-Switch Computing on Multi-GPUs are both tagged ISCA 2026. One entry does not line up with the rest: MoE Lens is labelled ICLR 2025 while carrying a March 2026 date, so either the venue tag refers to where it was published and the date to when it was added, or one of the two is wrong. ArXiv preprints dominate the count regardless of venue, which is the other thing the tags show: a large share of what is new in this area has not been through a review process yet.
The newest entries cluster around load balancing, kernels and quantization
Reading only the June 2026 block gives a fair impression of where the work is going, and it is systems work. UltraEP: Unleash MoE Training and Inference on Rack-Scale Nodes with Near-Optimal Load Balancing is dated 2 June. LoopMoE: Unifying Iterative Computation with Mixture-of-Experts for Language Modeling is dated 3 June. Less is MoE: Trimming Experts in Domain-Specialist Language Models heads the list at 4 June. Just below them come three ICML entries on structure and safety: PRISM on self-organized expert specialization for vision foundation models, DOT-MoE on differentiable optimal transport for what the authors call MoEfication, and DAG-MoE on moving from a simple mixture to structural aggregation. May brings the deployment cluster: an attention and feed-forward disaggregation design-space study for MoE serving, ReMoE on router fine-tuning under a memory constraint, GEMQ on global expert-level mixed-precision quantization, ROMER on expert replacement and router calibration for analog compute-in-memory, and MixServe as an automatic distributed serving system with hybrid parallelism and a fused communication algorithm. The disaggregation paper deserves its own note, because its title is a question rather than a system name: it frames attention and feed-forward disaggregation as a design space to be explored for efficient MoE serving, which is the shape a field takes when it has more placement options than it has settled conventions.
April and March bring specialization studies and training infrastructure
The middle of the list shifts from systems toward the question of what the experts actually learn, which is the part of the literature a survey is best placed to organize. Do Domain-specific Experts exist in MoE-based LLMs? asks the question directly on 7 April. The Expert Strikes Back: Interpreting Mixture-of-Experts Language Models at Expert Level tries to answer what an individual expert has become, on 2 April. DBES, a systematic benchmark and metric suite for evaluating expert specialization in large-scale MoEs, appears on 18 May, and Sparse Models, Sparse Safety: Unsafe Routes in Mixture-of-Experts LLMs on 9 February treats routing as a safety surface rather than an efficiency one. MESA, improving MoE safety alignment via decentralized expertise, is an ICML entry in the same vein. On the training side, UniEP is a unified expert-parallel MoE mega-kernel for LLM training, Scalable Training of Mixture-of-Experts Models with Megatron Core is the infrastructure counterpart, and Holistic Scaling Laws for Optimal Mixture-of-Experts Architecture Optimization tries to say what the right shape is rather than what the right router is. Two more entries push the same line of inquiry: Variational Routing, a scalable Bayesian framework for calibrated MoE transformers, tagged ICML and dated 10 March, which treats calibration as the problem rather than efficiency, and Expert Divergence Learning for MoE-based language models from 10 February, tagged ICLR.
Multimodal and non-language work sits alongside the language entries
A third of the visible list is not about language models at all, which is a useful corrective if you arrive expecting a language-only reading list. The multimodal entries include VidPrism, a heterogeneous mixture of experts for image-to-video transfer, tagged CVPR 2026; SMoES, soft modality-guided expert specialization in MoE vision-language models, also CVPR; On Token's Dilemma, dynamic MoE with drift-aware token assignment for continual learning of large vision language models, CVPR; and MoE-GRPO, optimizing mixture-of-experts through reinforcement learning in vision-language models, also CVPR. Computer vision has its own cluster, with a paper on the design and behaviour of sparse MoE layers in CNN-based semantic segmentation and another on enhancing specialization through cluster-aware upcycling, both CVPR entries from 15 April. Time series forecasting gets WaveMoE, a wavelet-enhanced MoE foundation model, tagged ICLR 2026. And applied work reaches medicine, with RANGER, sparsely-gated MoE with adaptive retrieval re-ranking for pathology report generation, tagged CVPR 2026. The breadth is the point: the same sparse routing idea keeps reappearing in domains that share nothing but a matrix.
Corrections go to an email address, and the badge asks for pull requests
The maintenance conventions are informal and worth knowing if you intend to help. Mistakes and suggestions are invited by email at an address on an hkust-gz.edu.cn subdomain rather than through the issue tracker, which is a fast way to fix a typo in a date and a slow way to argue about a taxonomy. The badges at the top of the page do the usual three things: a listing on awesome.re, a PRs welcome badge, and a last-commit badge wired to a repository name that is slightly different from the one you cloned, since the badge points at withinmiaov/A-Survey-on-Mixture-of-Experts while the repository is A-Survey-on-Mixture-of-Experts-in-LLMs. That is the kind of small inconsistency a list this long accumulates. The repository publishes no releases, its last push to the default branch was on 18 August 2026, and it is not archived. Its table of contents has four entries, a taxonomy, the paper list organized chronologically and categorically, contributors and a star history, which is the shape of a curated index rather than a software project.
Editorial conclusion
Use this repository as a map of the mixture-of-experts literature rather than as a text to read through. Its value is that it is dated and categorised, so you can see what changed in a given month, which venues are publishing MoE work, and which models are open. It will not tell you how any of the mechanisms work, because the content lives in the TKDE paper on arXiv. Three things to check before you rely on it. Whether an entry's venue tag and date agree with each other, since at least one entry is labelled with a 2025 conference and a 2026 date. Whether the survey version you read matches the list, since a January 2025 version is referenced and the repository has been pushed as recently as 18 August 2026. And whether a paper you need is open, because the entries above the arrow are the ones you can read the code for, and everything below it is described but not released.
Frequently asked questions
How does a mixture of experts work in LLMs?
This repository is the official page for a survey on the topic, accepted by TKDE and available on arXiv at 2407.06204, and it does not explain the mechanism itself. What the page does provide is a chronological, colour-coded index of MoE papers covering routing, load balancing, expert specialization, quantization and distributed training, with open-source models above an arrow and proprietary ones below it.
What is A-Survey-on-Mixture-of-Experts-in-LLMs?
The official GitHub page for the survey paper A Survey on Mixture of Experts in Large Language Models, accepted by TKDE. The repository itself holds only a README, a licence and an assets folder, and the README is the paper list rather than the paper.
How does the MoE survey repository organise its paper list?
Chronologically by model release date, with a venue tag and an explicit date on each entry. A timeline places open-source models above an arrow and proprietary ones below it, and four colours mark the domains: NLP in green, computer vision in yellow, multimodal in pink and recommender systems in cyan.
Where can I read the MoE survey paper itself?
On arXiv at 2407.06204, which is where the README points for the detailed contents rather than summarising them. The README also notes a previous version dated January 2025.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/withinmiaov-a-survey-on-mixture-of-experts-in-llms)