SageMaker Python SDK v3: the unified ModelTrainer and ModelBuilder API
A library for training and deploying machine learning models on Amazon SageMaker
At a glance
- What is it?
- The SageMaker Python SDK is AWS's library for training and deploying models on Amazon SageMaker. Version 3 splits it into four packages and replaces Estimator, Model and Predictor with ModelTrainer and ModelBuilder, so existing v2 code will not run unchanged.
- Who is it for?
- Adopt v3 for new SageMaker projects that start from the v3-examples notebooks, and stay on 2.x if your codebase imports Estimator, Model or Predictor, because those interfaces are not supported in 3.x. Before migrating, check the migration.md notes and confirm that sagemaker-core, sagemaker-train, sagemaker-serve and sagemaker-mlops resolve against your pinned environment, since the top-level sagemaker package is now only a wrapper around those four.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What the SageMaker Python SDK is for
The library wraps the SageMaker control plane in Python classes so you do not hand-assemble API calls for every training job and endpoint. It targets developers who already have training data in S3 and want to launch a job on managed instances, or who have a model artifact and want a hosted endpoint behind it. The README frames the scope as training and deploying models with Apache MXNet, PyTorch, Amazon algorithms, or your own algorithm packaged in a SageMaker-compatible Docker container. That last case matters most: the SDK is the glue between your container and SageMaker's orchestration, not a training framework of its own. If you are not on AWS, or if you are running training on your own Kubernetes cluster, nothing here applies. The SDK assumes IAM roles, S3 paths and SageMaker instance types as the vocabulary of every call.
Why v3 replaced Estimator, Model and Predictor
Version 3.0.0 is a breaking release, and the README is explicit that older interfaces such as Estimator, Model and Predictor and all their subclasses are not supported in V3. The replacement is two unified classes. ModelTrainer replaces the Estimator class and the framework-specific classes such as PyTorchEstimator and SKLearnEstimator. ModelBuilder replaces the Model class and framework-specific model classes such as PyTorchModel, TensorFlowModel, SKLearnModel and XGBoostModel. The stated reasoning is reduced boilerplate: instead of picking the right framework class, you construct one trainer or one builder and pass the container image or model artifact. The README also describes the API as object-oriented with auto-generated configs aligned with AWS APIs. That is a real shift in how much of the request shape the SDK decides for you. Where v2 asked for instance_count, instance_type and output_path as flat arguments, v3 moves input channels into a separate InputData object with channel_name and data_source. More objects, fewer positional arguments.
How the v3 packages are laid out
The repository root holds four directories next to the tests and docs: sagemaker-core, sagemaker-train, sagemaker-serve and sagemaker-mlops. Each is published separately on PyPI under the same name. The top-level sagemaker distribution in pyproject.toml declares dependencies on sagemaker-core>=2.21.0,<3.0.0, sagemaker-train>=1.21.0,<2.0.0, sagemaker-serve>=1.21.0,<2.0.0 and sagemaker-mlops>=1.21.0,<2.0.0, and its own [tool.setuptools] section declares packages = [], so installing sagemaker pulls in the four subpackages rather than shipping a large module tree itself. Optional extras exist for the three workload packages: train, serve, mlops, and all. Note the version ranges are asymmetric: core is capped below 3.0.0 while the other three are capped below 2.0.0. The Python requirement is >=3.10, and the classifiers list 3.10, 3.11 and 3.12. The project classifies its own development status as Alpha, which is worth reading literally when you pin it in a production image.
Installing the SageMaker Python SDK and running a first training job
The README gives the upgrade path directly: pip install --upgrade sagemaker moves you to the latest 3.x. If you need the older interfaces, the documented downgrade is pip install sagemaker==2.*. Both commands are reproduced from the README's migration section.
pip install --upgrade sagemakerAfter that, the first real use is a training job. The README's v3 example imports ModelTrainer from sagemaker.train and InputData from sagemaker.train.configs, constructs the trainer with a training_image and a role, wraps the S3 prefix in an InputData with channel_name and data_source, and calls trainer.train with input_data_config as a list. The role string in the example is an IAM role ARN, and the data source is an S3 URI; neither is optional in practice.
from sagemaker.train import ModelTrainer
from sagemaker.train.configs import InputData
trainer = ModelTrainer(
training_image="my-training-image",
role="arn:aws:iam::123456789012:role/SageMakerRole"
)
train_data = InputData(
channel_name="training",
data_source="s3://my-bucket/train"
)
trainer.train(input_data_config=[train_data])For inference, the README's counterpart example builds a ModelBuilder with a model name and a model_path pointing at an S3 object, then calls build() and invoke() on the returned endpoint. The README writes the invoke argument as an ellipsis, so the payload shape is not specified there; check the inference examples in v3-examples before assuming a format.
from sagemaker.serve import ModelBuilder
model_builder = ModelBuilder(
model="my-model",
model_path="s3://my-bucket/model.tar.gz"
)
endpoint = model_builder.build()
result = endpoint.invoke(...)The repository also carries a v3-examples directory with notebooks for custom distributed training, distributed local training, hyperparameter training, JumpStart training, local training, HuggingFace inference, in-process mode and inference specs. Those notebooks are the practical starting point, because the README's inline snippets stop at the constructor.
Migrating from v2 is a rewrite, not a version bump
The README states plainly that Estimator, Model, Predictor and all their subclasses are not supported in V3. There is no compatibility shim described. The repository does include a migration.md file at the root, but the README itself does not walk through a per-class mapping beyond the two training and inference examples. The conversion in those examples is not mechanical: the v2 Estimator took instance_count, instance_type and output_path as constructor arguments, while the v3 ModelTrainer example passes only training_image and role and moves the data channel out into InputData. That means the instance configuration has to move somewhere else, and the README does not show where in the snippet. Anyone with a v2 training script should expect to rewrite the job configuration, not just the import lines. The downgrade path exists and is documented, which is the honest answer for a codebase that cannot absorb the rewrite on a deadline.
Limits and cases where this SDK is the wrong layer
The SDK only makes sense on SageMaker. If your training runs on EC2, EKS or on-premises, the classes here give you nothing, and you would be paying for an abstraction over a service you are not using. The second limit is version coupling. The top-level package pins sagemaker-core below 3.0.0 and the other three below 2.0.0, so a release of any subpackage outside those ranges will not resolve through the umbrella package until the constraints are updated. If you install sagemaker-train directly to get a newer version, you are deliberately stepping outside the combination the umbrella package was tested against. Third, the classifiers mark the project as Development Status 3 - Alpha even though it is the official AWS library; treat the API surface as still moving between minor releases. Finally, if all you need is to start a training job from a script, the AWS CLI or boto3 will do it with fewer layers, and the SDK's value only appears when you are managing many jobs and endpoints with shared configuration.
SageMaker Python SDK compared with boto3
boto3 is the general AWS SDK for Python and exposes the SageMaker API as raw request and response dictionaries. The SageMaker Python SDK sits above that surface with opinionated classes. The difference shows up in the training call: with boto3 you assemble the CreateTrainingJob parameters yourself, including the algorithm specification, input data config and output data config, and you poll DescribeTrainingJob for status. With the v3 SDK you construct a ModelTrainer and call train, and the SDK builds the request from the object graph. The trade-off is control. boto3 accepts any parameter the SageMaker API accepts, including ones the SDK has not modeled yet, and it never breaks your code on a major version bump of an abstraction layer. The SDK gives you shorter code and the v3-examples notebooks as a starting pattern. If you are debugging why a job launched with the wrong container or the wrong channel name, boto3's explicit dictionaries are easier to inspect than an object that generates configs for you.
Maintenance, licensing and what to verify before adopting
The repository is not archived, and the last push was on 2026-09-10, with releases v3.21.0 on 2026-08-25, v3.20.0 on 2026-08-14 and v3.19.0 on 2026-08-11. That release cadence means the four subpackages move quickly, and the pin ranges in pyproject.toml will lag behind individual subpackage releases at times. Budget for periodic dependency bumps rather than a set-and-forget pin. The licence is Apache-2.0, and the classifiers confirm OSI approval; that permits commercial and closed-source use, but you should confirm the terms of any SageMaker service you call separately, since the licence covers the library and not the AWS service. The README does not document a rollback procedure for a deployed endpoint, and it does not state a support window for the 2.x line beyond the pip install sagemaker==2.* command. Before committing, read migration.md at the repository root and open the v3-examples notebook closest to your workload, since the README's snippets stop before instance configuration and payload shapes.
Editorial conclusion
Adopt v3 for new SageMaker projects that start from the v3-examples notebooks, and stay on 2.x if your codebase imports Estimator, Model or Predictor, because those interfaces are not supported in 3.x. Before migrating, check the migration.md notes and confirm that sagemaker-core, sagemaker-train, sagemaker-serve and sagemaker-mlops resolve against your pinned environment, since the top-level sagemaker package is now only a wrapper around those four.
Frequently asked questions
How do I install the SageMaker Python SDK?
The README's migration section gives pip install --upgrade sagemaker for the latest 3.x. If you need the older Estimator, Model and Predictor interfaces, the documented alternative is pip install sagemaker==2.*.
What is the SageMaker Python SDK?
It is an open source library for training and deploying machine learning models on Amazon SageMaker. The README describes support for Apache MXNet, PyTorch, Amazon algorithms, and your own algorithms packaged in SageMaker-compatible Docker containers.
What is the difference between the SageMaker Python SDK and boto3?
The SDK provides classes such as ModelTrainer and ModelBuilder that build the SageMaker requests for you, while boto3 is the general AWS SDK for Python that exposes the underlying API directly. The SDK's stated goal is reduced boilerplate; boto3 gives you every API parameter without an abstraction layer.
What does the Python SDK do?
In this project's case it wraps SageMaker training and hosting so you can launch jobs and endpoints from Python instead of assembling API calls. The README's v3 examples show a ModelTrainer for training and a ModelBuilder for inference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/aws-sagemaker-python-sdk)