Library / SDK
deepseek-ai/DeepJIT avatar
deepseek-ai/DeepJIT

DeepJIT: Header-Only C++20 JIT Runtime for CUDA and Ascend Kernels

A lightweight library for xPU kernel JIT compilation

349 stars33 forksC++License varies

At a glance

What is it?
DeepJIT is a header-only C++20 library from DeepSeek that gives CUDA and Huawei Ascend extension authors a shared interface for compiling kernel source at runtime, caching the resulting binaries, and launching them with backend-specific options. It is designed to be embedded into a kernel library, not used as a standalone application.
Who is it for?
C++ and Python extension authors who need to compile and cache CUDA kernels at runtime and who want the same interface to also target Huawei Ascend NPUs will find DeepJIT useful: add include/ to the target's include path, include the backend header, and create a lazy runtime. The shared cache support via DJ_JIT_CACHE_DIR is immediately practical in cluster environments where multiple processes or nodes can reuse the same compiled artifacts.
Can I use it commercially?
Not without permission. GitHub finds no licence file in the repository, and without a licence all rights are reserved by default: you may read the code but not reuse it. Check the README, or ask the authors, before using it.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly C++, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

Editorial analysis

What Problem DeepJIT Solves

Writing a custom CUDA kernel library that also runs on Huawei Ascend NPUs requires two separate JIT compilation stacks: NVCC and the CUDA Driver API for NVIDIA hardware, and Bisheng with ACL for Ascend. If both backends share logic for caching compiled binaries, hashing source code and include files, handling lazy initialization, and integrating with PyTorch's stream model, that shared logic must be written twice or maintained in a way that does not diverge.

DeepJIT provides one interface for both. A consumer library includes the CUDA or Ascend backend header, calls create_lazy_jit with a configuration object, and gets a runtime object that handles compilation, caching, and launch. The backend-specific details (NVCC flags and CUBIN format for CUDA; Bisheng, ld.lld, and ACL for Ascend) stay inside the library. The consumer sees the same compile/load/launch workflow for both targets.

The library is header-only and intended to be embedded into an extension's source tree as a vendored dependency, not installed system-wide.

Architecture: Shared Runtime, Backend-Specific Compilers

The repository layout reflects the separation between shared and backend-specific code. The include/deep_jit/runtime/ directory holds the configuration, cache key computation, and lazy initialization machinery that both backends use. The include/deep_jit/backend/cuda/ and include/deep_jit/backend/ascend/ directories each contain their own compiler integration, device query code, kernel loader, and launch option types. The include/deep_jit/cache/ directory holds the in-memory and on-disk cache implementations shared by both backends. The include/deep_jit/python_api.hpp file provides a pybind11 registration helper for consumer libraries that expose their runtime to Python.

Cache keys account for kernel source, tracked include files, compiler versions, effective compiler options, and an application-provided dependency signature. Two compiled kernels with different source but identical other inputs get different cache entries; two compilations of the same kernel with the same inputs reuse the same entry.

The tests/ directory contains CUDA and Ascend integration tests along with example extensions under tests/test_cuda_proj/ and tests/test_ascend_proj/ that demonstrate how to wire DeepJIT into a real project.

Integrating DeepJIT into a C++ Extension

Add the include directory to the CMake target's include path:

cmake
target_include_directories(my_target PRIVATE third-party/deep_jit/include)

For a CUDA backend, include the CUDA entry header:

cpp
#include <deep_jit/backend/cuda/backend.hpp>

For Ascend, include the Ascend entry header instead:

cpp
#include <deep_jit/backend/ascend/backend.hpp>

Create a lazy runtime at the library's global scope. Lazy initialization defers device and compiler discovery until the runtime is first used:

cpp
inline auto jit = deep_jit::create_lazy_jit<deep_jit::CUDA>(
    deep_jit::Config("/absolute/path/to/my_library", "MYLIB"));

For the Ascend backend, use deep_jit::Runtime<deep_jit::Ascend> in place of deep_jit::CUDA. The README notes that if configuration is only known during library initialization, start with an empty lazy object and assign its factory later.

The CUDA backend requires CUDA headers version 12.4 or higher and NVCC version 12.9 or higher. The Ascend backend requires CANN with the bin/bisheng and bin/ld.lld executables, the Ascend adv_api headers, and ACL and torch_npu headers and runtime. Both backends require Linux, a C++20 compiler and standard library with std::format support, Python, and pybind11.

The CUDA backend also exposes compilation controls for inspecting the output of the compilation process. The README documents the ability to dump CUDA PTX and SASS output from compiled kernels, which is useful for performance profiling and debugging. The CUDA backend additionally supports a Python post-compilation hook that runs after each successful kernel compilation. The Ascend backend provides its own assembly inspection capability. Both backends support runtime defaults and per-kernel overrides for compiler options, so individual kernels can use different optimization flags without changing the global configuration.

Shared Kernel Cache Across Processes and Nodes

The on-disk cache uses a directory structure with atomic rename for publishing completed entries. The README states that concurrent processes can compile the same kernel and one will reuse the other's published result once the rename completes. Multiple users and nodes can point to the same cache directory over a distributed filesystem that provides atomic directory rename and POSIX fsync semantics.

Configure the cache directory with an environment variable:

bash
export DJ_JIT_CACHE_DIR=/shared/deep_jit

For a personal writable cache combined with a read-only shared lookup cache:

bash
export DJ_JIT_CACHE_DIR="$HOME/.dj:/shared/deep_jit"

DeepJIT searches cache roots in order and writes misses only to the first root. The README notes that the shared lookup cache can be read-only for consumers who should not write to it. The README also recommends using a trusted shared directory, since the cache contains compiled binaries.

The README documents a cache warmup feature for pre-compiling kernels based on historical usage patterns as in development and not yet available.

Limitations and Deployment Constraints

The repository lists no license in its metadata. The README gives no explicit permission for commercial use. This is a significant constraint for any team that wants to embed DeepJIT in a product and needs a clear open-source license before doing so.

The Ascend backend targets Huawei's CANN environment. CANN is a proprietary SDK distributed by Huawei and is not widely available outside Huawei hardware deployments. Teams without access to Ascend hardware or CANN cannot use the Ascend backend, making the library effectively CUDA-only for most users.

Two features are documented as in development and not yet available: cache warmup from history and the Python compilation API for passing kernel source directly from Python. The README marks both with WIP labels.

AMD ROCm and Intel GPUs are not supported. Teams that need JIT kernel compilation across AMD hardware must look elsewhere. PyTorch's torch.utils.cpp_extension provides a similar JIT compilation interface for CUDA extensions but operates at a different level: it compiles C++ extension modules, not raw CUDA kernel source strings. DeepJIT targets the lower-level case of compiling and launching individual kernels from source at runtime. A kernel library that manages a collection of specialized CUDA kernels and needs to compile the right variant at inference time, based on a data-dependent shape or hardware configuration discovered at runtime, is the use case DeepJIT is designed for. A library that only needs to compile a fixed set of kernels at build time has no need for JIT at all and should use the static NVCC build path instead.

Maintenance and License Status

The last push to this repository was on 2026-09-28. The repository is not archived and shows active development. There are no GitHub releases; the main branch carries the current state of the library. The main authors listed in the README are @guyan364, @kurisu6912, and @LyricZhao. The repository does not include a LICENSE file and the license field in the repository metadata is listed as unknown. The README does not state any terms of use. Before using DeepJIT in a commercial or distributed product, confirm the licensing terms with the project authors directly by opening a GitHub issue. The README documents the expected API surface and the CMakeLists.txt at the root is described as for debugging and IDE indexing rather than for building the library as a standalone artifact.

Editorial conclusion

C++ and Python extension authors who need to compile and cache CUDA kernels at runtime and who want the same interface to also target Huawei Ascend NPUs will find DeepJIT useful: add include/ to the target's include path, include the backend header, and create a lazy runtime. The shared cache support via DJ_JIT_CACHE_DIR is immediately practical in cluster environments where multiple processes or nodes can reuse the same compiled artifacts. Two features documented as in development (cache warmup from history and the Python compilation API) are not yet available and should not be treated as present functionality. The repository lists no license, which is a material consideration for commercial adoption: check with the project authors before embedding DeepJIT in a product.

Frequently asked questions

Does DeepJIT work with PyTorch?

Yes. The include/deep_jit/python_api.hpp file provides pybind11 registration so a consumer library can expose its configured runtime to Python. The CUDA backend uses the current PyTorch CUDA stream by default, and the Ascend backend uses torch_npu streams.

What GPU hardware does DeepJIT support?

DeepJIT supports NVIDIA CUDA GPUs through NVCC and the CUDA Driver API, and Huawei Ascend NPUs through Bisheng, ld.lld, and ACL. AMD ROCm and Intel GPU backends are not documented in the repository.

How does DeepJIT handle duplicate kernel compilations across multiple processes?

The on-disk cache uses atomic directory rename to publish completed entries. Concurrent processes can compile the same kernel and the published result is reused by the process that arrives second. A shared cache directory configured via DJ_JIT_CACHE_DIR lets multiple nodes on a distributed filesystem share compiled artifacts.

Official sources

  1. deepseek-ai/DeepJIT on GitHub
  2. Issues
  3. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/deepseek-ai-deepjit.svg)](https://hysenlabs.com/projects/deepseek-ai-deepjit)