Taichi Lang: GPU Kernels Written in Python
Productive, portable, and performant GPU programming in Python.
At a glance
- What is it?
- Taichi Lang embeds a JIT-compiled parallel language inside Python, so a decorated Python function becomes CUDA, Vulkan or CPU machine code. It suits simulation and numerical code that has outgrown NumPy but does not justify hand-written CUDA.
- Who is it for?
- Adopt Taichi Lang if your workload is a stencil, particle or field simulation that you want to express in Python and run on a GPU without writing CUDA. Do not adopt it as a drop-in NumPy replacement for array algebra or as a general deep learning framework; the README positions it around kernels, fields and SNodes, and the differentiable programming and quantized computation features are documented as separate, partly experimental paths.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 85 days ago.
- What is it written in?
- Mainly C++, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap Taichi Lang fills between NumPy and hand-written CUDA
NumPy gives you array expressions, but it gives you no way to write an explicit per-element loop that stays fast. Once your algorithm needs a while loop per particle, a sparse grid, or per-element branching, you either vectorize it into something unreadable or move to CUDA. Taichi Lang takes the second path without the second language. The README describes it as "an open-source, imperative, parallel programming language for high-performance numerical computation" that is embedded in Python and uses just-in-time compiler frameworks such as LLVM to offload compute-intensive Python code to native GPU or CPU instructions. The audience named in the repository metadata is scientific and research code: real-time physical simulation, numerical computation, visual effects, robotics and vision. If your inner loop is a for loop over pixels or particles rather than a matrix multiply, this is the case it was built for.
Kernels, fields and SNodes: the actual execution model
Three constructs carry the design. A function marked with @ti.kernel is compiled by the JIT compiler into parallel machine code; a function marked with @ti.func is inlined into a kernel rather than called from Python. Data lives in ti.field objects with a dtype and a shape, and the README's fractal example shows the loop idiom: writing for i, j in pixels iterates over every element of a 2D field, and the comment in the example states that this is parallelized over all pixels. There is no explicit thread index and no block size to choose. The second construct is the SNode, which the README calls "a set of generic data containers" for composing hierarchical, multi-dimensional fields, and which it links to spatially sparse computing. That hierarchy is what lets a field be dense in one region and sparse in another, which is the mechanism behind the sparse-computation topic in the repository metadata. Portability comes from the backend list: CUDA, Vulkan, OpenGL 4.3+, Apple Metal, x64 and ARM CPUs, with WebAssembly marked experimental. The same kernel source targets whichever backend ti.init selects.
Installing Taichi Lang and running a first kernel
The README gives two commands to get started. The first installs the released package, and the second launches a demo gallery that ships with it, which is the fastest way to confirm that a backend actually works on your machine before you write anything.
pip install taichi # Install Taichi Lang
ti gallery # Launch demo galleryThe README also notes a nightly package hosted on the project's own index, with an explicit warning that nightly packages "may crash because they are not fully tested" and that validity is not guaranteed. Treat that channel as opt-in for testing unreleased features.
pip install -i https://pypi.taichi.graphics/simple/ taichi-nightlyFor a first kernel of your own, the README's fractal example is the canonical shape: initialize with ti.init(arch=ti.gpu), allocate a field, decorate a kernel, and loop over the field's indices. The loop body in that example runs a Julia set iteration and writes a float back into pixels[i, j]. A window titled "Julia Set" opens through ti.GUI, and the animation updates as the kernel is called repeatedly with a changing time value. If Taichi Lang is installed correctly, the README states you should see that animation.
import taichi as ti
ti.init(arch=ti.gpu)
n = 320
pixels = ti.field(dtype=float, shape=(n * 2, n))
@ti.kernel
def paint(t: float):
for i, j in pixels:
pixels[i, j] = 1 - t * 0.02The prerequisites are worth reading before you start, because they are narrow. The README lists Windows, Linux and macOS, Python 3.6 through 3.10 in 64-bit builds only, and the backend options above. The setup.py classifiers list Python 3.9, 3.10 and 3.11, which does not match the README's stated range, so verify your interpreter version against the documentation rather than assuming the classifiers are the newer truth.
Where Taichi Lang is the wrong tool
Two boundaries stand out. First, the README states Python 3.6 to 3.10, 64-bit only. If your environment is pinned to a newer interpreter, or to a 32-bit target, the documented support does not cover you, and the mismatch with the setup.py classifiers means you cannot resolve the question from the metadata alone. Second, the nightly channel is explicitly untested; anyone who needs reproducible builds should stay on the released versions, and the release cadence visible in the tags is uneven, with v1.7.3 in December 2024 and v1.7.4 in July 2025. Beyond versions, the model itself is a constraint: Taichi Lang is built around kernels and fields, not around tensor algebra. If your work is matrix multiplication, convolution and autograd over dense arrays, a tensor framework is the shorter path, and the README's own framing of Taichi as a language for numerical computation rather than a tensor library supports that reading. The differentiable programming support is documented as its own topic rather than as the default mode, and quantized computation is labelled experimental in the README. Neither is something to build a production pipeline on without checking the linked documentation first.
How Taichi Lang differs from Numba and from CuPy
The closest comparison is Numba, which also compiles decorated Python functions to machine code. The difference is what each assumes about your data. Numba compiles functions that operate on NumPy arrays passed in from outside; the arrays are ordinary host memory and the compiler reasons about them. Taichi Lang instead gives you its own field abstraction, with an SNode hierarchy underneath that can be dense or spatially sparse, and the kernel loops over that field directly. That hierarchy is the reason sparse computation appears as a topic in the repository metadata: a sparse grid is expressed as a field layout choice rather than as index bookkeeping in your own code. CuPy is a different split again. It reimplements the NumPy API on the GPU, so existing array expressions move over largely unchanged, but you are still writing array expressions. Taichi Lang asks you to write loops. The trade is that loops over particles, stencils and adaptive grids are natural in Taichi Lang and awkward in an array API, while plain linear algebra is natural in CuPy and more verbose in Taichi Lang. The README also notes integration with NumPy and PyTorch, so the two approaches are not mutually exclusive within one project.
Licence, build cost and what a version upgrade involves
Taichi Lang is Apache-2.0, and the setup.py classifier confirms the Apache Software License. That is a permissive licence with an explicit patent grant, which matters for a compiler that will be embedded in other products. It is not legal advice, and the LICENSE file at the repository root is the text that governs. Upgrading the released package is a pip install, but the cost sits underneath: the package is a compiled C++ extension, and pyproject.toml lists setuptools, wheel, numpy, pybind11, cmake, scikit-build and ninja as build requirements. Building from source therefore pulls in a full CMake and pybind11 toolchain, and setup.py documents optional environment variables for that path, including TAICHI_CMAKE_ARGS for extra CMake arguments and build-type variables such as DEBUG, RELWITHDEBINFO and MINSIZEREL. The README points anyone who needs experimental features or a custom environment at the developer installation page rather than at pip. Practically, that means binary wheels are the normal upgrade route and source builds are a deliberate project. The last push to the default branch was on 2026-07-06, so the repository is not dormant, but the gap between the v1.7.3 and v1.7.4 tags in the release list is a reminder to check the release notes rather than assume a steady stream of changes.
Editorial conclusion
Adopt Taichi Lang if your workload is a stencil, particle or field simulation that you want to express in Python and run on a GPU without writing CUDA. Do not adopt it as a drop-in NumPy replacement for array algebra or as a general deep learning framework; the README positions it around kernels, fields and SNodes, and the differentiable programming and quantized computation features are documented as separate, partly experimental paths. Before committing, verify that your Python version falls inside the 3.6 to 3.10 range the README lists, check that your target backend (CUDA, Vulkan, OpenGL 4.3+, Metal, or the experimental WebAssembly) is actually available on your machine, and run ti gallery to confirm the toolchain works before porting an existing solver.
Frequently asked questions
How do I install Taichi Lang in Python?
The README gives a single pip command, pip install taichi, and notes that the package supports Windows, Linux and macOS on 64-bit Python 3.6 through 3.10. A separate nightly package is available from the project's self-hosted index, but the README warns it may crash because it is not fully tested.
What is Taichi Lang used for?
The README describes it as an imperative, parallel programming language for high-performance numerical computation, embedded in Python and JIT-compiled through frameworks such as LLVM. It names real-time physical simulation, numerical computation, visual effects, robotics and vision among its applications.
How do I install Taichi Lang?
The README's installation section uses pip install --upgrade taichi, followed by ti gallery to launch the bundled demo gallery. Building from source is a separate path documented on the developer installation page, and the README directs people there for experimental features or custom environments.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/taichi-dev-taichi)