TimeCraft: Microsoft Research's Diffusion Framework for Time Series Generation
Official code for TimeCraft: A Time Series Generation Framework for Real-World Applications
At a glance
- What is it?
- TimeCraft is a Python framework from Microsoft Research that generates synthetic time series using diffusion models, with built-in support for cross-domain generalization, text-based control over output patterns, and influence-function-guided adaptation toward downstream task performance. It bundles five related sub-projects in a single repository.
- Who is it for?
- TimeCraft is the right tool for research teams working on time series generation who need a framework that goes beyond single-domain replication. The repository is structured as a collection of sub-projects, each with its own code directory, so adoption means picking the specific component relevant to your problem: TimeDP for cross-domain work, CaTSG for causal constraints, OATS for augmenting foundation model pretraining, or Diff-MN for irregular observations.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 54 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Problem TimeCraft Addresses
Most existing time series generation methods train and generate within a single domain. A model trained on medical signals does not transfer to financial time series without full retraining, because the statistical patterns differ enough that the learned distribution does not generalize. The TimeCraft README frames this as a fundamental limitation of prior work.
A second gap is controllability. Generating data unconditionally, without guiding the output toward specific trends, seasonality, or domain characteristics, limits how useful the synthetic data is in practice. A forecasting team that needs data resembling a rare demand spike cannot get that from an unconditional generator.
A third gap is that many methods optimize for distributional similarity to training data rather than for improving downstream task performance. Synthetic data that looks statistically similar to real data may not help a downstream classifier or forecaster any more than simply using the original training data.
TimeCraft's design addresses all three of these gaps, though the approach to each is documented in separate sub-projects within the repository rather than in a single unified pipeline.
Three Core Technical Mechanisms
Cross-domain generalization in TimeCraft comes from a set of semantic prototypes, described in the README as analogous to a dictionary of temporal patterns. These prototypes encode domain-invariant features such as trends and seasonality. A Prototype Assignment Module, referred to as PAM in the README, computes domain-specific weights for these prototypes from a small number of examples in the target domain. The result is a domain prompt, a latent representation that captures the target domain without requiring retraining on it.
Text-based control works through a multi-agent text generation system that produces textual descriptions of time series patterns. These descriptions are used to build paired time series and text data during training. At inference, a text prompt guides the generation through a hybrid framework that combines semantic prototypes with free-form text input. The README notes that text carries semantic information and domain knowledge that can guide generation in a more interpretable way than purely numerical conditioning.
Target-aware adaptation uses influence functions to measure the expected reduction in task-specific loss that each synthetic sample would contribute. Instead of generating data that mimics the training distribution, this mechanism generates data that is specifically shaped to improve performance on a downstream task such as forecasting, classification, or anomaly detection. The README describes this as influence-guided diffusion.
Repository Layout and Sub-Projects
The repository is not a single installable library. It is a collection of related research code directories, each corresponding to a published or submitted paper. Understanding the layout is necessary before attempting to use any part of it.
The top-level directories are: BRIDGE, CaTSG, DiGA, Diff-MN, OATS, TarDiff, TimeDP, demo_dev, diffusion, and process. The core train and inference entry point is train_inference.py at the top level. Configuration is managed through environment.yml, which specifies the conda environment.
CaTSG adds causal constraints to the diffusion process, generating time series that adhere to specified causal structures rather than just statistical patterns. The README describes this as enabling robust what-if analysis. OATS focuses on online data augmentation for time series foundation model pretraining, synthesizing model-tailored samples during training to improve zero-shot performance. Diff-MN addresses irregularly sampled observations, modeling continuous latent dynamics to generate high-resolution outputs from sparse inputs.
The supplementary directory and figures directory support reproducibility for the published results. Teams who want to run experiments from a specific paper should read the README in the corresponding sub-project directory.
Setting Up the Environment
TimeCraft uses a conda environment defined in environment.yml at the repository root. The README does not provide step-by-step install instructions in the main text; it points to the wiki and the sub-project directories for detailed setup. The environment file is the starting point for dependency management.
The process directory and demo_dev directory appear to hold helper scripts for data preparation and demonstration runs. The train_inference.py entry point at the root handles training and inference across the framework, though the specific arguments it accepts are not reproduced in the README itself.
Because the repository has no GitHub releases and no versioned packages on PyPI, there is no pinnable release to install. Teams adopting TimeCraft must work directly from the repository at a specific commit. The README does not specify minimum Python or CUDA version requirements in the main text.
Genuine Constraints and Cases Where TimeCraft Is the Wrong Tool
TimeCraft is research code, not a production library. The repository has no releases, no changelog, and no documented upgrade path. This is common for academic code and is not a criticism of the work itself, but it is important to state plainly for teams evaluating it for operational use.
The multi-agent text generation system requires a language model to produce textual descriptions of time series. The README does not specify which language model is used, what API access it requires, or what the inference cost looks like at scale. Teams that want text-based control cannot plan infrastructure costs from the README alone.
Cross-domain generalization depends on having a small number of few-shot examples from the target domain. The README describes PAM as using few-shot examples to compute domain-specific prototype weights. Teams working with entirely novel domains where even a handful of representative samples are unavailable may not be able to use the cross-domain capability.
The influence function approach in target-aware adaptation has known computational cost at scale. Influence functions require computing per-sample gradients, which is expensive for large datasets or large models. The README acknowledges the approach but does not quantify the overhead.
How TimeCraft Compares to TimeGAN
TimeGAN is an earlier approach to time series generation that uses a generative adversarial network architecture. It is a commonly cited baseline in the time series generation literature and represents the prior generation of methods that TimeCraft is designed to improve on.
The key difference in approach is architecture: TimeGAN uses a GAN, which requires balancing generator and discriminator training and is known to be sensitive to hyperparameter choices. TimeCraft uses diffusion models, which train by learning to reverse a gradual noising process. The README does not provide a direct benchmark comparison between the two, but the motivation section positions TimeCraft's diffusion foundation as addressing the controllability and generalization limitations of earlier methods.
TimeCraft also explicitly targets multiple domains in a single model, while TimeGAN is trained per-domain. For a team that needs to generate data for several domains without maintaining separate models for each, that architectural choice is a practical advantage.
Maintenance and Research Context
The repository is not archived and its last push was on 2026-08-07. It is published under the MIT license and maintained under the Microsoft organization on GitHub. The README links to three Microsoft Research blog posts about TimeCraft and TimeDP, which provide narrative context for the technical choices.
The project has no GitHub releases. The News and Updates section of the README documents three 2026 additions: CaTSG, OATS, and Diff-MN, each with paper links and code directories. This suggests the repository is actively extended as new research results are published rather than maintained as a stable library.
The MIT license permits unrestricted use including in commercial products. The repository includes a CODE_OF_CONDUCT.md and a SECURITY.md, which are standard Microsoft open-source repository files. A citation or attribution expectation is not documented in the README, though teams publishing work that builds on TimeCraft would typically cite the original papers linked in the README.
Editorial conclusion
TimeCraft is the right tool for research teams working on time series generation who need a framework that goes beyond single-domain replication. The repository is structured as a collection of sub-projects, each with its own code directory, so adoption means picking the specific component relevant to your problem: TimeDP for cross-domain work, CaTSG for causal constraints, OATS for augmenting foundation model pretraining, or Diff-MN for irregular observations. Teams that need a production-grade, maintained library with versioned releases and documented upgrade paths will not find that here: the repository has no GitHub releases. Researchers who want to reproduce or build on the published results in the referenced papers will find the code organized to match those papers.
Frequently asked questions
What Python dependencies does TimeCraft require?
TimeCraft provides an environment.yml file at the repository root for setting up a conda environment. The README does not list specific dependency versions in the main text; the environment file is the authoritative source. There are no versioned releases on PyPI.
Can TimeCraft generate time series without any text input?
TimeCraft supports three flexible input branches: domain-only (semantic prototypes), text-only, and a hybrid of both. The README states that users can activate any one, any two, or all three inputs depending on the application, so text input is optional.
Is TimeCraft suitable for real-time or production data generation?
TimeCraft is research code without versioned releases or a documented deployment path. The README frames it as a framework for real-world applications in the sense of research benchmarks and experimental settings, not as production infrastructure. Teams needing production-grade synthesis pipelines will need to build that infrastructure around the research code themselves.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/microsoft-timecraft)