HYSEN LABS DIRECTORY

NVIDIA

Verified repositories, classifications and analysis from the NVIDIA organization.

44 curated open-source projects
★ 14,743

TensorRT LLM provides users with an easy-to-use Python API to define Large Language Models (LLMs) and supports state-of-the-art optimizations to perform inference efficiently on NVIDIA GPUs. TensorRT LLM also contains components to create Python and C++ runtimes that orchestrate the inference execution in a performant way.

Python2,781 forks
NVIDIA

TensorRT

★ 13,374

NVIDIA® TensorRT™ is an SDK for high-performance deep learning inference on NVIDIA GPUs. This repository contains the open source components of TensorRT.

C++2,411 forks
NVIDIA

cosmos

★ 11,943

NVIDIA Cosmos is an open platform of world models, datasets, and tools that enables developers to build Physical AI for robots, autonomous vehicles, smart infrastructure, and more.

Jupyter Notebook894 forks
NVIDIA

cutlass

★ 10,508

CUDA Templates and Python DSLs for High-Performance Linear Algebra

C++2,107 forks
NVIDIA

OpenShell

★ 10,338

OpenShell is the safe, private runtime for autonomous AI agents.

Rust1,400 forks
★ 9,668

Samples for CUDA Developers which demonstrates features in CUDA Toolkit

C++2,433 forks
NVIDIA

garak

★ 9,368

the LLM vulnerability scanner

Python1,312 forks
NVIDIA

apex

★ 9,001

A PyTorch Extension: Tools for easy mixed precision and distributed training in Pytorch

Python1,530 forks
★ 8,146

NVIDIA Isaac GR00T N1.7 - A Foundation Model for Generalist Robots.

Python1,480 forks
NVIDIA

DALI

★ 5,769

A GPU-accelerated library containing highly optimized building blocks and an execution engine for data processing to accelerate deep learning training and inference applications.

C++680 forks
NVIDIA

cuml

★ 5,281

NVIDIA cuML: GPU-Accelerated Machine Learning

Python677 forks
NVIDIA

nccl

★ 5,098

Optimized primitives for collective multi-GPU communication

C++1,416 forks
★ 4,684

A unified library of SOTA model optimization techniques like quantization, distillation, pruning, neural architecture search, speculative decoding, etc. It compresses deep learning models for downstream deployment frameworks like TensorRT-LLM, TensorRT, vLLM, etc. to optimize inference speed.

PythonAI & Machine Learning663 forks
★ 4,197

Generative AI reference workflows optimized for accelerated infrastructure and microservice architecture.

Jupyter Notebook1,101 forks