Sana
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
HYSEN LABS DIRECTORY
Verified repositories, classifications and analysis from the NVlabs organization.
SANA: Efficient High-Resolution Image Synthesis with Linear Diffusion Transformer
Lightning fast C++/CUDA neural network framework
VILA is a family of state-of-the-art vision language models (VLMs) for diverse multimodal AI tasks across the edge, data center, and cloud.
Welcome to GR00T Whole-Body Control (WBC)! This is a unified platform for developing and deploying advanced humanoid controllers. This includes: Decoupled WBC models used in NVIDIA Isaac-Gr00t, Gr00t N1.5 and N1.6 and GEAR-SONIC
Eagle: Frontier Vision-Language Models with Data-Centric Strategies
cuda-oxide is a Rust-to-CUDA compiler that lets you write (SIMT) GPU kernels in safe(ish), idiomatic Rust. It compiles standard Rust code directly to PTX — no DSLs, no foreign language bindings, just Rust.
[CVPR 2024 Highlight] FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
SoL-Pi: Scaling Auto-Research Loops for Efficient Agent Harnesses
Long Video Gen Infrastructure
Official code for the CVPR 2022 (oral) paper "Extracting Triangular 3D Models, Materials, and Lighting From Images".
[CVPR 2025] Official PyTorch Implementation of MambaVision: A Hybrid Mamba-Transformer Vision Backbone
NVIDIA Alpamayo 1 Nano is an open 10B reasoning VLA model for autonomous vehicles that pairs driving trajectories with Chain-of-Causation reasoning.
Sionna: An Open-Source Library for Research on Communication Systems
AlpaSim is an open-source autonomous vehicle simulation platform designed for development and testing of end-to-end AV policies
Kernel Design Agents (KDA) is a agent-centric workflow to write high-performance CUDA Kernels.
Official implementation of "Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding"
[ICML2024 (Oral)] Official PyTorch implementation of DoRA: Weight-Decomposed Low-Rank Adaptation
ToolOrchestra is an end-to-end RL training framework for orchestrating tools and agentic workflows.
OmniVinci is an omni-modal LLM for joint understanding of vision, audio, and language.
[ICLR 2026]QeRL enables RL for 32B LLMs on a single H100 GPU.