# Amazon SageMaker Examples: Official AWS Jupyter Notebooks for Machine Learning

> aws/amazon-sagemaker-examples is the official AWS repository of self-contained Jupyter notebooks covering every major phase of building a machine learning model on Amazon SageMaker: training, model customization, evaluation, inference, and MLOps. It is a learning and reference library, not a runnable application, and every notebook requires an active AWS account to execute.

**aws/amazon-sagemaker-examples** — Example 📓 Jupyter notebooks that demonstrate how to build, train, and deploy machine learning models using 🧠 Amazon SageMaker. 

- Repository: https://github.com/aws/amazon-sagemaker-examples
- Website: https://sagemaker-examples.readthedocs.io
- Stars: 10,991 · Forks: 6,956
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/aws-amazon-sagemaker-examples

## What amazon-sagemaker-examples Contains and Who Uses It

The repository holds official Jupyter notebooks that demonstrate the features of Amazon SageMaker. Every notebook is self-contained: it keeps its scripts, data files, and images in its own folder, and all file references are relative to that folder. The design means you can clone the repository, navigate to any notebook, open it in JupyterLab, and run it without resolving cross-notebook dependencies.

The primary audience is ML practitioners already building on AWS who want a concrete, working example of a specific SageMaker capability rather than reading documentation alone. The notebooks also serve as integration tests for the SageMaker team: they demonstrate that features work end to end, not just that the API exists.

Pure learners who want to understand machine learning concepts first will find the notebooks assume familiarity with the AWS console, IAM policies, and S3. The notebooks do not teach ML fundamentals or Python basics. They start from the assumption that you already have a trained model or training script and want to understand how SageMaker's managed infrastructure handles a particular workflow.

The repository is distinct from the SageMaker Example Community repository (aws/amazon-sagemaker-examples-community), which accepts contributions beyond official feature coverage. The official repository, maintained directly by the Amazon SageMaker team, accepts only notebooks that demonstrate a SageMaker feature not yet covered elsewhere in the repository.

## Five Notebook Categories and How They Map to a Model's Lifecycle

The repository organizes notebooks into five top-level directories that follow the lifecycle of a typical ML project.

The training/ directory contains notebooks for Amazon SageMaker Training, a managed service that containerizes workloads and handles compute provisioning. Notebooks cover script and framework training, distributed training across multiple instances, managed spot training with checkpointing, heterogeneous clusters, bringing your own container, and the @remote decorator that runs local code as a SageMaker training job. The newer examples use ModelTrainer from SageMaker Python SDK v3.

The model_customization/ directory addresses fine-tuning pre-trained foundation models. The six techniques covered are supervised fine-tuning (SFT), direct preference optimization (DPO), reinforcement learning with verifiable rewards (RLVR), reinforcement learning from AI feedback (RLAIF), multi-turn reinforcement learning (MTRL), and continued pre-training (CPT). The notebooks also show recipe overrides, data mixing, JumpStart fine-tuning with a private model hub, and distributed fine-tuning across serverless, serverful, and HyperPod compute.

The evaluation/ directory shows how to score customized models using the SageMaker evaluator surface. Notebooks demonstrate standard benchmark evaluation, custom scoring functions, an LLM-as-judge pattern, and Inspect AI. The goal of each notebook is to compare a base model and a fine-tuned model on the same footing so the improvement is measurable.

The inference/ directory covers deployment options and scaling. Notebooks use ModelBuilder and sagemaker-core as the deployment surface. The README does not fully enumerate the inference patterns in the visible portion, but SageMaker's inference documentation covers real-time endpoints, batch transform, serverless inference, and asynchronous inference.

The mlops/ directory covers operationalizing models at scale, including pipeline orchestration, experiment tracking, and model registry patterns.

## Setup Requirements Before Running Any Notebook

The README is direct about what is needed before any notebook can run. Four things are required: an AWS account, proper IAM User and Role setup, an Amazon SageMaker Notebook Instance (or SageMaker Studio), and an S3 bucket.

The AWS account requirement is a hard gate. There is no local-only mode for these notebooks. Every training job, every endpoint deployment, and every model evaluation call reaches out to AWS APIs and incurs compute charges. The IAM Role must have the permissions SageMaker needs to call other AWS services on your behalf; the documentation linked in the README covers the specifics of what the role requires.

The quickest path described in the README is to use an Amazon SageMaker Notebook Instance. When you open a Notebook Instance in the AWS console, the example notebooks are automatically loaded and accessible through the SageMaker Examples tab in Jupyter or the SageMaker logo in JupyterLab. That one-click access means you do not need to clone the repository manually if you are already working inside a Notebook Instance.

Running notebooks outside of a SageMaker Notebook Instance is possible but requires updating the IAM role definition in each notebook and installing the necessary libraries locally. The README describes this as requiring minimal modification, but the actual scope of modification depends on how each notebook initializes the SageMaker session and role. Notebooks that rely on the automatic role detection inside a Notebook Instance will need explicit role ARN configuration when run from a developer's laptop or a CI environment.

The S3 bucket stores training data, model artifacts, and outputs. Notebook code that writes to S3 typically uses a default bucket name derived from the account and region, which the SageMaker SDK creates if it does not exist. Teams with strict S3 naming policies or bucket-level restrictions will need to configure the bucket name explicitly before running.

## SageMaker-Core and SDK v3 in the v1.0.0 Release

The README opens with an announcement about SageMaker-Core, described as a new Python SDK that provides an object-oriented interface for SageMaker resources. SageMaker-Core introduces resource chaining, which lets developers pass resource objects (a TrainingJob object, a Model object) as parameters to subsequent calls rather than passing string ARNs and names manually. The stated goal is to eliminate the manual parameter plumbing that Boto3 requires.

The v1.0.0 release, tagged on 2026-09-08 as SageMaker Python SDK v3 Golden Example Notebooks, marks the point where the official example notebooks align with SDK v3 and its SageMaker-Core API. The training/ notebooks use ModelTrainer from SDK v3; the inference/ notebooks use ModelBuilder and sagemaker-core. This means notebooks from before v1.0.0, still present in the repository for context, use the older SDK v2 API surface.

For teams migrating from SDK v2 to SDK v3, the repository now has side-by-side reference material showing the older and newer patterns. The SageMaker-Core documentation is linked from the README at sagemaker.readthedocs.io.

SageMaker-Core also adds auto code completion and type hints, which the README cites as a developer experience improvement over Boto3's dynamically typed API. This is relevant for teams writing production ML pipelines, where type errors in job configuration parameters historically surfaced only at runtime, not at development time.

## What the Repository Does Not Cover

The repository covers SageMaker features as they stand at the time each notebook was written. It does not cover the AWS Bedrock API, which is a separate service for foundation model inference. Notebooks that call a foundation model through SageMaker JumpStart are in scope; notebooks that call Bedrock directly are not, and there is no Bedrock directory in the repository layout.

The README states that the repository is focused entirely on covering SageMaker feature breadth. It does not contain reference architectures for specific industries (healthcare, financial services, retail), end-to-end application templates, or notebooks that demonstrate integrating SageMaker with non-AWS databases and services. Those use cases go to the community repository instead.

Notebooks that demonstrate a feature already covered by an existing notebook in the repository are not accepted. The README explicitly tells contributors to check for existing coverage before submitting a pull request, since a new notebook demonstrating, for example, a second way to do distributed training with the same framework will be rejected in favor of the existing example. This policy keeps the repository focused but means some alternative approaches to the same problem are not documented here.

The Makefile in the repository root is for building the Sphinx documentation site (sagemaker-examples.readthedocs.io), not for running notebooks. The index.rst, conf.py, and related .rst files are documentation source files for the site, not notebook orchestration scripts.

## The Community Repository as the Overflow Channel

Amazon maintains a second repository, aws/amazon-sagemaker-examples-community, for examples that fall outside the official scope. The README describes it as containing additional examples and reference solutions beyond the official repository. This community repository accepts contributions from engineers and solution architects at AWS, not only from the core SageMaker team.

The practical implication is that if you have a working notebook that demonstrates a valid SageMaker workflow but shows a feature already covered in the official repository, the community repository is the right submission target. The official repository's contribution guidelines redirect these contributions explicitly.

For readers, this means the official repository and the community repository are complementary. The official repository has the canonical, maintained examples of each core SageMaker feature. The community repository has additional patterns, alternative approaches, industry-specific solutions, and experiments that did not fit the official coverage mandate.

The split also affects maintenance expectations. Official repository notebooks are maintained by the SageMaker team and updated when the SDK changes; community notebooks are maintained by the contributor who submitted them or by AWS employees with time to spare. When running a community notebook, check the last commit date to assess whether it reflects current SDK behavior.

## Maintenance, Contribution Policy, and the Apache-2.0 License

The repository is licensed under Apache-2.0. That license allows commercial use, modification, and distribution. Derivative works that distribute modified versions must include the Apache-2.0 notice and any modifications must be documented. For internal enterprise use, Apache-2.0 places no distribution obligations as long as you are not distributing the modified code externally.

The last push to the default branch was on 2026-09-09, and v1.0.0 shipped on 2026-09-08. That is recent, which indicates the repository is actively maintained. The release note for v1.0.0 calls it the golden examples baseline for SageMaker Python SDK v3, suggesting this is a deliberate checkpoint after substantial SDK changes rather than a routine minor release.

Contribution is governed by the policy in CONTRIBUTING.md. The key constraint is that new notebooks must cover a SageMaker feature not yet addressed anywhere in the repository. Pull requests that duplicate existing coverage are not accepted; the README directs those to the community repository. CODEOWNERS in the repository root defines who reviews each area.

The environment.yml in the repository root defines a Conda environment for contributors building the documentation locally. The tox.ini and conf.py are configuration for the Sphinx documentation pipeline. These files are for documentation contributors; notebook contributors work in their own self-contained notebook folder and do not need the documentation toolchain.

For teams concerned about long-term availability, the Apache-2.0 license and the public repository mean the notebooks can be forked and maintained internally if AWS ever discontinues the official repository.

## Conclusion

The amazon-sagemaker-examples repository is the right starting point for ML practitioners who are committed to the AWS ecosystem and want working, runnable examples of SageMaker features. It is not useful to teams evaluating whether to use AWS at all, because every notebook requires an active account, IAM permissions, and an S3 bucket before it can run. Teams contributing new notebooks should check whether their feature is already covered before opening a pull request; the repository's own README states that coverage is the primary acceptance criterion. The last push was on 2026-09-09, and v1.0.0 shipped on 2026-09-08 as the golden examples baseline for the SageMaker Python SDK v3.

## FAQ

### What is Amazon SageMaker and what is it used for?

Amazon SageMaker is a managed AWS service for building, training, and deploying machine learning models. The amazon-sagemaker-examples repository provides official Jupyter notebooks demonstrating each major SageMaker capability, from managed training jobs to real-time inference endpoints.

### Is Amazon SageMaker free?

The amazon-sagemaker-examples notebooks themselves are free to use under Apache-2.0, but running them incurs AWS compute charges for training jobs, endpoints, and the Notebook Instance or Studio session. An active AWS account with billing enabled is required before any notebook can execute.

### What are some common use cases for Amazon SageMaker?

The repository covers training models on managed infrastructure, customizing foundation models through fine-tuning techniques such as SFT and DPO, evaluating model quality with automated scoring, deploying models to real-time and batch inference endpoints, and building MLOps pipelines for production model management.

### Is SageMaker better than Azure ML?

The amazon-sagemaker-examples repository does not compare SageMaker to Azure ML. The repository covers SageMaker-specific features and the SageMaker Python SDK; any comparison to other managed ML platforms is outside its scope.

## Sources

- [aws/amazon-sagemaker-examples on GitHub](https://github.com/aws/amazon-sagemaker-examples)
- [License: Apache-2.0](https://github.com/aws/amazon-sagemaker-examples/blob/default/LICENSE)
- [Project website](https://sagemaker-examples.readthedocs.io)
- [README](https://github.com/aws/amazon-sagemaker-examples/blob/default/README.md)
- [Releases](https://github.com/aws/amazon-sagemaker-examples/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/aws-amazon-sagemaker-examples
