Library / SDK
microsoft/tensorwatch avatar
microsoft/tensorwatch

TensorWatch: Real-Time ML Training Debugging in Jupyter Notebooks

Debugging, monitoring and visualization for Python Machine Learning and Data Science

3,472 stars361 forksJupyter NotebookMIT

At a glance

What is it?
TensorWatch is a Microsoft Research debugging and visualization tool for Python machine learning training that streams metrics to Jupyter Notebooks in real time and supports live queries against a running training process, at the cost of significant security constraints on its deployment environment.
Who is it for?
TensorWatch is appropriate for individual researchers running ML training on a private machine who want real-time stream visualization and interactive queries in Jupyter Notebooks. It is not appropriate for shared servers, multi-tenant environments, or any setting where the ZMQ port or training machine might be reachable by untrusted users, because its Lazy Logging feature executes arbitrary Python expressions from clients.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Jupyter Notebook, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 30, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What TensorWatch Does and Who It Is For

Machine learning training runs produce a stream of scalar values: loss curves, accuracy metrics, learning rates, gradient norms. Tracking these in real time requires either a logging service like TensorBoard or custom plotting code that runs alongside the training loop. TensorWatch provides a third option: stream these values to a Jupyter Notebook where you can visualize them live and run arbitrary Python queries against the training process while it is running.

TensorWatch comes from Microsoft Research and targets data scientists and deep learning researchers working in Jupyter Notebooks. The README states it supports Python 3.x and was tested with PyTorch 0.4 through the 1.x series, and that most features work with TensorFlow eager tensors as well.

The last push to the repository was on 2026-03-30. The project is at version 0.9.1. The license is MIT.

Installation and First Stream

Install TensorWatch with pip:

code
pip install tensorwatch

TensorWatch uses graphviz for network architecture diagrams. On some platforms this requires a separate manual installation of the graphviz system package before the Python bindings work.

Here is the minimal usage pattern from the README. This code logs integer pairs to a stream once per second and generates a Jupyter Notebook to view them:

python
import tensorwatch as tw
import time

# streams will be stored in test.log file
w = tw.Watcher(filename='test.log')

# create a stream for logging
s = w.create_stream(name='metric1')

# generate Jupyter Notebook to view real-time streams
w.make_notebook()

for i in range(1000):
    # write x,y pair we want to log
    s.write((i, i*i))

    time.sleep(1)

After running this script, open the generated test.ipynb in Jupyter Notebook, choose Cell > Run all, and the line graph updates in real time as new values arrive. Streams are persisted in the test.log file, so they can be replayed after the training run completes.

Lazy Logging Mode: Live Queries Against Training Processes

The most distinctive TensorWatch feature is what the README calls Lazy Logging Mode. Instead of deciding in advance which metrics to log, a WatcherClient can send a Python expression string to the Watcher server at any time during a training run. The server evaluates that expression using Python's eval() in the context of the running process and returns the result as a stream. This lets you ask questions about the model's internal state while training is happening, not just track values you thought to log beforehand.

This design has direct security consequences documented clearly in the SECURITY.md file. Any client that can connect to the Watcher's ZMQ port and holds the HMAC key can execute arbitrary Python code in the training process. By default, the Watcher binds to 127.0.0.1 (localhost only) and messages are HMAC-SHA256 signed, but the eval() execution is fundamental to the feature and cannot be removed without disabling it entirely.

Security Architecture and Mitigations

The README documents four security risks explicitly. The first is the eval() execution path in Lazy Logging (CWE-94). The second is pickle deserialization of ZMQ messages (CWE-502); mitigated by HMAC verification before deserialization and a RestrictedUnpickler allowlist of permitted modules (builtins, collections, numpy, torch, pandas, tensorwatch). The third is pickle deserialization when reading saved stream files; the same allowlist applies. The fourth is YAML loading in the bundled hiddenlayer utilities, mitigated by using yaml.SafeLoader by default.

For multi-process setups where the Watcher and WatcherClient run in separate OS processes, the HMAC key must be set explicitly before calling initialize():

bash
export TENSORWATCH_HMAC_KEY=$(python -c "import os; print(os.urandom(32).hex())")

The README's guidance is explicit: do not expose TensorWatch ports to untrusted networks or users, do not run it on machines where untrusted users have local network access, and treat TensorWatch data files with the same caution as executable scripts. This is not a hardened security tool; it is a development-only debugging tool with documented mitigations that assume a trusted user environment.

Visualization Capabilities and Custom Streams

TensorWatch supports multiple built-in visualizer types for different metric types, displayed inline in Jupyter Notebook via ipywidgets. The architecture separates what gets logged (a stream of Python objects) from how it is displayed (the visualizer). Because streams contain arbitrary Python objects, a visualizer can consume anything: scalar values, images, text, model outputs, or structured data.

The README describes this as going beyond the traditional approach where what you log is exactly what you see. The Lazy Logging path allows querying values that were never explicitly logged: internal layer activations, gradient magnitudes, parameter distributions, or any other expression valid in the training process's scope.

For offline analysis, streams persist to .log files using pickle. These can be loaded after training and explored through the same Jupyter interface, subject to the same RestrictedUnpickler constraint on what module types are allowed to be deserialized.

Limitations: Version Currency and Production Unsuitability

TensorWatch was built and tested against PyTorch 0.4 through the 1.x series. The PyTorch API has changed considerably in subsequent major versions, and the README does not state compatibility with more recent releases. Teams using current PyTorch versions should treat compatibility as something to verify before adopting TensorWatch.

The tool is explicitly not designed for production environments. The SECURITY.md carries a prominent warning that TensorWatch is a development and debugging tool and is not suitable for multi-tenant or adversarial environments. This is not a cautionary caveat to be ignored; the eval() execution path means a malicious user with ZMQ access can run arbitrary code in the training process.

The dependency on Jupyter Notebook also bounds where TensorWatch is useful. Training jobs that run on remote compute clusters or in containers without Jupyter server access require additional setup to use TensorWatch's interactive features. The file-based stream persistence works in those environments, but the real-time interactive query capability requires network connectivity to the Jupyter kernel.

Alternative: TensorBoard and the Trade-off in Approach

TensorBoard, included with TensorFlow and usable with PyTorch via the tensorboard package, is the standard alternative for ML training visualization. TensorBoard uses a file-based event log: training code writes events to a directory, and TensorBoard serves a web interface that reads those files. This log-first approach has no eval() execution path and poses no equivalent security risk.

The trade-off is flexibility versus safety. TensorBoard requires deciding in advance what to log; adding new metrics means code changes and a new training run. TensorWatch's Lazy Logging Mode allows post-hoc queries without restarting training, which is genuinely useful during exploratory research where the interesting questions emerge while the model is training. The cost is the single-user, trusted-environment constraint that makes TensorWatch unsuitable anywhere TensorBoard would be deployed safely.

Editorial conclusion

TensorWatch is appropriate for individual researchers running ML training on a private machine who want real-time stream visualization and interactive queries in Jupyter Notebooks. It is not appropriate for shared servers, multi-tenant environments, or any setting where the ZMQ port or training machine might be reachable by untrusted users, because its Lazy Logging feature executes arbitrary Python expressions from clients. Before deploying it, read the SECURITY.md file and set the TENSORWATCH_HMAC_KEY environment variable in any multi-process setup to ensure HMAC authentication is active across processes.

Frequently asked questions

Is TensorWatch safe to use on a shared server?

No. The README and SECURITY.md state explicitly that TensorWatch is not designed for production, multi-tenant, or adversarial environments. Its Lazy Logging feature executes arbitrary Python expressions from clients over ZMQ. It should only run on a machine where all local network users are trusted.

Does TensorWatch work with TensorFlow as well as PyTorch?

The README states that most TensorWatch features work with TensorFlow eager tensors. It was tested primarily with PyTorch versions 0.4 through 1.x. The TensorFlow compatibility is described as partial rather than full.

Can TensorWatch stream data from a remote training machine?

By default, the Watcher binds to 127.0.0.1 (localhost only). Streaming from a remote machine requires explicitly changing the binding to expose the ZMQ port over the network. The SECURITY.md warns against this in any environment with untrusted users, since doing so exposes the eval() execution path.

Official sources

  1. Issues
  2. License: MIT
  3. microsoft/tensorwatch on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/microsoft-tensorwatch.svg)](https://hysenlabs.com/projects/microsoft-tensorwatch)