Scalene: a Python CPU, GPU and memory profiler with AI optimization proposals
Scalene: a high-performance, high-precision CPU, GPU, and memory profiler for Python with AI-powered optimization proposals
At a glance
- What is it?
- Scalene profiles CPU, GPU and memory use in Python and can ask an AI provider for rewrite suggestions. This article covers how to install it, how the profiler separates Python from native time, and where it stops being the right tool.
- Who is it for?
- Adopt Scalene if you need line-level CPU, GPU and memory attribution for a Python program and you are willing to read a report rather than a flame graph. Do not adopt it if your bottleneck lives in a compiled extension you cannot edit, or if sending source lines to an external AI provider is not acceptable for your codebase.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What Scalene measures that a sampling profiler does not
Most Python profilers answer one question: which function consumed the most time? Scalene answers a narrower and more awkward set of questions. It reports CPU time, GPU time and memory allocation per line, and it separates time spent in Python code from time spent in native code such as C extensions and library calls. That split matters because the fix is different in each case. A hot Python loop can be rewritten or moved to NumPy. A hot native call usually cannot, and the useful response is to call it less often or with better arguments.
The author list is Emery Berger, Sam Stern and Juan Altmayer Pizzorno, and the project is developed at UMass Amherst under the plasma-umass organization. The README claims Scalene runs orders of magnitude faster than many other profilers while producing more detailed output, and it describes the tool as the first profiler to incorporate AI-powered proposed optimizations. The intended audience is developers and researchers who already have a slow Python program and need to know which line to change, not people looking for a general application performance monitor.
How Scalene attributes time to a line
The mechanism is a mix of instrumentation and sampling, not pure sampling. The repository ships C++ sources under src/source, including libscalene.cpp and get_line_atomic.cpp, plus vendored Heap-Layers and printf directories that the Makefile clones during a source build. That native layer is what lets Scalene intercept allocations and attribute them to a line rather than to a function. A pure Python tracer cannot see inside a C extension; a native interposer can.
The command structure changed to a verb-based form. There are two commands, run and view. Profiling writes a JSON profile, and viewing renders it in a browser, in the terminal, or as HTML. That separation is the important architectural detail: the profile is a file, so you can collect it on one machine and inspect it on another, or keep it as an artifact. The README's own example shows the default output filename as scalene-profile.json and notes that view --standalone produces a self-contained HTML file.
Memory attribution is per allocation site, which is why the profiler can point at a specific line that allocates a large object rather than at the function that contains it. GPU profiling depends on nvidia-ml-py, declared in pyproject.toml with the marker platform_system !='Darwin', so GPU metrics are not available on macOS through that path.
Installing Scalene and running a first profile
Installation is a normal Python package install from PyPI. The README gives this as the first option, and conda-forge as the second. Use a virtual environment so the native extension is built against the interpreter you intend to profile.
python3 -m pip install -U scaleneAfter installation, profile a script by passing it to the run subcommand. The README states that this saves the profile to scalene-profile.json by default.
scalene run your_prog.pyTo pass arguments through to your own program, the README uses a triple-hyphen separator rather than the double hyphen you may expect.
scalene run your_prog.py --- --arg1 --arg2Then open the result. Running view with no arguments opens the profile in a browser; adding --cli renders it in the terminal instead.
scalene view
scalene view --cliIf you only care about CPU and want a faster run, the README lists --cpu-only. You can also name the output file explicitly.
scalene run --cpu-only your_prog.py
scalene run -o results.json your_prog.pyOptions can be stored in a YAML file and loaded with -c or --config, which the README illustrates with scalene run -c scalene.yaml your_prog.py. The truncated README example begins with an output key, so treat the config schema as something to confirm against scalene run --help-advanced rather than assume.
The AI optimization proposals and what they send off your machine
This is the feature that gets Scalene cited in IEEE Spectrum and discussed on the Real Python podcast, and it deserves a plain description. According to the README, you select an AI provider in the box under AI Optimization Options and enter credentials if the provider needs them. Supported providers listed there are Amazon Bedrock, Microsoft Azure, OpenAI, and local models via Ollama. Once configured, clicking the lightning bolt beside a line or the explosion icon for a region generates a proposed optimization, and clicking a proposal copies it to the clipboard. Repeated clicks generate different suggestions.
The README is candid that results vary, saying your mileage may vary while also claiming that in some cases the suggestions are impressive. Two things follow from that. First, this is a suggestion engine, not a verified transformation; nothing in the documented flow applies a patch or re-runs the profile to confirm the change helped. Second, unless you choose the Ollama path, generating a proposal means sending the relevant code to a third-party service. The README documents provider selection and credential entry but does not describe a redaction step, so the decision to use a remote provider is a decision about your source code leaving your machine.
Where Scalene is the wrong tool
Scalene profiles a single process. The README's usage examples all invoke a script directly, and there is no documented mode for attaching to an already-running server, sampling a fleet of workers, or reconstructing a request path across services. If your production problem is latency spread across a queue, a database and three microservices, a Python line profiler is the wrong instrument.
Interpreter support is another boundary. pyproject.toml declares requires-python = ">=3.8,!=3.11.0", so Python 3.11.0 specifically is excluded even though 3.11 generally is not. The classifiers list 3.8 through 3.14. If you are pinned to the exact 3.11.0 release, pip will refuse the install, and the fix is to move to a later 3.11 patch rather than to work around the pin.
There is also a cost to the detail. Profiling every allocation and separating native from Python time is more work than sampling a call stack, and the README's ordering advice for the VS Code extension notes that your code should run for at least a second for a profile to appear. Short scripts and microbenchmarks that finish in milliseconds are a poor fit. The README does not document a rollback or undo path for AI proposals, because proposals are copied to the clipboard rather than applied.
Scalene compared with cProfile and py-spy
cProfile is in the standard library and reports call counts and cumulative time per function. It does not report memory allocation, does not separate Python from native time, and its per-call instrumentation is heavy enough that the measurement distorts the program. Scalene's pitch is the opposite trade: a native layer that costs more to build but attributes cost to lines and to allocation sites.
py-spy is the closest alternative in spirit. It samples from outside the process, which means it can attach to a running program without modifying it and without a rebuild. Scalene's documented usage is the inverse: you launch your program through scalene run, and the profiler is present from the start. If you need to inspect a process you cannot restart, an external sampler is the better fit. If you need to know which line allocated 400 MB, an external sampler cannot tell you, and that is the gap Scalene is built to fill.
Maintenance, releases and licence
The repository is not archived, and the last push was on 2026-08-27. Recent releases are v2.3.0 on 2026-05-12, described as new visualizations, accuracy improvements and bug fixes; v2.2.1 on 2026-03-22; and v2.1.4 on 2026-02-15, a maintenance release with bugfixes and performance improvements. That cadence suggests a project that ships fixes rather than one that has stalled.
Upgrade cost is mostly about the native extension. The repository has both a setup.py and a pyproject.toml, and the Makefile shows vendored dependencies (Heap-Layers and printf) being cloned at build time on the Windows path. If you install a wheel, you avoid that build. If you install from source, you inherit it, and a compiler problem becomes your problem. The declared dependencies include numpy, psutil, rich, Jinja2 and nvidia-ml-py, so a Scalene install pulls a real dependency set into your environment.
The licence is Apache-2.0, and the classifiers confirm OSI approval. Apache-2.0 is permissive and includes an explicit patent grant, which is generally friendlier to corporate adoption than a bare MIT licence, but the AI features raise a separate question the licence does not answer: what your chosen provider does with the code you send it. That is governed by the provider's terms, not by Scalene's licence. This is not legal advice; check both documents against your own policy.
Editorial conclusion
Adopt Scalene if you need line-level CPU, GPU and memory attribution for a Python program and you are willing to read a report rather than a flame graph. Do not adopt it if your bottleneck lives in a compiled extension you cannot edit, or if sending source lines to an external AI provider is not acceptable for your codebase. Before trusting a number, verify that the version you installed supports your interpreter: pyproject.toml declares requires-python >=3.8,!=3.11.0, so Python 3.11.0 itself is excluded.
Frequently asked questions
How do you use Scalene to profile a Python script?
Install it with python3 -m pip install -U scalene, then run scalene run your_prog.py. The README states the profile is saved to scalene-profile.json, and you open it with scalene view or scalene view --cli for the terminal.
How do you pass command line arguments to the program Scalene is profiling?
The README uses a triple-hyphen separator: scalene run your_prog.py --- --arg1 --arg2. Arguments after the separator go to your program rather than to Scalene.
Does Scalene profile GPU usage?
The README describes Scalene as a CPU, GPU and memory profiler. GPU support depends on the nvidia-ml-py package, which pyproject.toml declares only when platform_system is not Darwin, so that path is not available on macOS.
Which Python versions does Scalene support?
pyproject.toml declares requires-python = ">=3.8,!=3.11.0" and the classifiers list 3.8 through 3.14. Python 3.11.0 itself is excluded, so use a later 3.11 patch release.
Can Scalene suggest how to make my code faster?
Yes. The README describes AI-powered optimization proposals: you select a provider such as Amazon Bedrock, Microsoft Azure, OpenAI, or a local model via Ollama, then click the lightning bolt beside a line or the explosion icon for a region. The README notes results vary and that proposals are copied to the clipboard rather than applied.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/plasma-umass-scalene)