★ 3,537A library for accelerating Transformer models on NVIDIA GPUs, including using 8-bit and 4-bit floating point (FP8 and FP4) precision on Hopper, Ada and Blackwell GPUs, to provide better performance with lower memory utilization in both training and inference.
Python827 forks
★ 3,480Project brief: Agent Skills for NVIDIA products, install into Claude Code, Codex, and other coding agents to run Physical AI, robotics, simulation, CUDA, and RAG workflows end to end. Skill Catalog Product | Description | Skills | AIQ | NVIDIA AI-Q Blueprint - deploy local AI-Q services and run shallow or deep research workflows as agent skills.
Python420 forks
★ 3,389CUDA Python: Performance meets Productivity
Cython331 forks
★ 3,298Open-source deep-learning framework for building, training, and fine-tuning deep learning models using state-of-the-art Physics-ML methods
Python799 forks
★ 3,292Pytorch implementation of FlowNet 2.0: Evolution of Optical Flow Estimation with Deep Networks
Python752 forks
★ 2,980NeMo Retriever Library is a scalable, performance-oriented document content and metadata extraction microservice. NeMo Retriever Library uses specialized NVIDIA NIM microservices to find, contextualize, and extract text, tables, charts and images that you can use in downstream generative applications.
Python349 forks
★ 2,884NVIDIA GPU Operator creates, configures, and manages GPUs in Kubernetes
Go553 forks
★ 2,650The NVIDIA NeMo Agent toolkit is an open-source library for efficiently connecting and optimizing teams of AI agents.
Python773 forks
★ 2,523CUDA Core Compute Libraries
C++507 forks
★ 2,443`std::execution`, the standard C++ framework for asynchronous and parallel programming.
C++274 forks
★ 1,528Router that virtually distributes inference across connected devices in the home.
Go258 forks
★ 1,221LLM KV cache compression made easy
Python182 forks
★ 1,145Open-source deep-learning framework for exploring, building and deploying AI weather/climate workflows.
PythonAI & Machine Learning263 forks
★ 1,042RAFT contains fundamental widely-used algorithms and primitives for machine learning and information retrieval. The algorithms are CUDA-accelerated and form building blocks for more easily writing high performance applications.
Cuda252 forks
★ 953cuDNN Frontend is NVIDIA's modern, open-source entry point to the cuDNN library and a growing collection of high-performance open-source kernels.
Python299 forks
★ 857cuVS - a library for vector search and clustering on the GPU
Cuda238 forks
★ 707Differentiable signal processing on the sphere for PyTorch
Jupyter Notebook74 forks
★ 592NVIDIA NVSHMEM is a parallel programming interface for NVIDIA GPUs based on OpenSHMEM. NVSHMEM can significantly reduce multi-process communication and coordination overheads by allowing programmers to perform one-sided communication from within CUDA kernels and on CUDA streams.
C++112 forks
★ 570High-performance, light-weight C++ LLM and VLM Inference Software for Physical AI
Python134 forks
★ 550Our inference and training framework to run on the Cosmos Models
Python152 forks