Library / SDK
cvxpy/cvxpylayers avatar
cvxpy/cvxpylayers

cvxpylayers: differentiating a convex program, if it happens to be DPP

Differentiable convex optimization layers

2,141 stars195 forksPythonApache-2.0

At a glance

What is it?
CVXPYlayers lets PyTorch, JAX and MLX treat a CVXPY problem as a layer that solves in the forward pass and differentiates in the backward pass, using diffcp to solve the KKT system rather than unrolling a solver. The gate is DPP: every example asserts the problem is disciplined-parametrized, and the GPU path needs either Moreau, whose licence terms you are told to check, or a Julia stack with one dependency installed from git main.
Who is it for?
Use cvxpylayers if you have a convex optimization problem embedded in a learning pipeline and the problem satisfies disciplined parametrized programming, because the backward pass is a KKT linear solve rather than a finite-difference approximation, and DPP gives a well-defined derivative where an unrolled solver would not have one.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 11 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 5, 2026, and from our analysis. They are not legal advice.

Editorial analysis

DPP is a hard gate, and every example asserts it

The single most important constraint on using this library is one line in the Usage section, and it is repeated as an assertion in every code example.

Note that the parametrized convex optimization problems must be constructed in CVXPY, using DPP. And then, in the PyTorch example:

python
import cvxpy as cp
import torch

from cvxpylayers.torch import CvxpyLayer

n, m = 2, 3
x = cp.Variable(n)
A = cp.Parameter((m, n))
b = cp.Parameter(m)
constraints = [x >= 0]
objective = cp.Minimize(0.5 * cp.pnorm(A @ x - b, p=1))
problem = cp.Problem(objective, constraints)
assert problem.is_dpp()

layer = CvxpyLayer(problem, parameters=[A, b], variables=[x])
A_tch = torch.randn(m, n, requires_grad=True)
b_tch = torch.randn(m, requires_grad=True)

# solve the problem
(solution,) = layer(A_tch, b_tch)

# compute the gradient of the sum of the solution with respect to A, b
solution.sum().backward()

DPP is disciplined parametrized programming, a modelling discipline CVXPY enforces. It restricts how a parameter may appear in a problem: affine in parameters at most, and a small set of other well-behaved forms. The restriction exists because DPP is what makes the derivative of the solution tractable and unambiguous. A problem outside it may still be a perfectly good convex program that CVXPY solves happily, and it just cannot be differentiated by this method.

That is the gap to understand. The set of problems you can write in CVXPY is strictly larger than the set you can put in a CvxpyLayer, and the boundary is a modelling convention rather than a mathematical impossibility. So a user porting an existing CVXPY model into a learning pipeline can hit a wall that has nothing to do with convexity and everything to do with how they wrote it.

The failure mode is worth noting. The example asserts rather than raising with a message, so the practical experience of hitting the boundary is an AssertionError, possibly with assertion optimisation stripping it out of a production build. The library is not being unhelpful on purpose; it is demonstrating the precondition in the most compact way. But a user who does not read the Usage section gets a bare assertion instead of an explanation of DPP and a link to the tutorial.

The log-log convex path has the same structure with a stricter check. For geometric and log-log convex programs the example asserts problem.is_dgp(dpp=True), which asks both that the problem is a disciplined geometric program and that it is additionally disciplined-parametrized. A model that is a valid geometric program but not DPP still fails.

The workaround, if you hit the boundary, is to rewrite the parameterisation so the parameters enter affinely. That is sometimes a small change and sometimes a substantial reformulation, and whether it is worth doing depends entirely on whether the rest of your pipeline needs the convex structure to be inside the autodiff graph.

The backward pass is a KKT linear solve, not solver autograd

The mechanism is stated in three sentences at the top of the README, and understanding them is what separates this from a wrapper.

A convex optimization layer solves a parametrized convex optimization problem in the forward pass to produce a solution. It computes the derivative of the solution with respect to the parameters in the backward pass.

What that second sentence means in practice is the interesting part, and the dependency list answers it. The core dependencies are CVXPY at 1.9.0 or newer, diffcp at 1.1.0 or newer, NumPy at 1.22.4 or newer, and SciPy at 1.13.0 or newer. diffcp is from the same organisation as CVXPY, and it is the component that does the differentiation.

The distinction is between two ways to differentiate through an optimisation. The first is to unroll the solver: treat every iteration of the interior point method as an operation in the graph and let autodiff propagate through all of them. That gives you a gradient, but it is the gradient of the approximate iterate rather than of the true solution, it costs memory proportional to the number of iterations, and it inherits whatever conditioning the solver had. The second is to use the fact that a convex program's solution is characterised by the KKT conditions, and to differentiate those conditions as a linear system. That is what diffcp does, and the derivative it computes is the exact derivative of the solution map for a conic program.

So the derivative is not an approximation of a gradient. It is the gradient. That is the paper's contribution, and it is why the DPP restriction and the KKT approach are the same design decision rather than two separate ones: DPP is the discipline under which the derivative of the solution map is a well-defined function of the parameters, and the KKT system is how you compute it.

SciPy is in that list because the KKT system is a large sparse linear system and solving it at scale needs a sparse factorisation. That is also the answer to a question the README does not address directly, which is what the per-backward-pass cost actually is. You are not running an optimiser in the backward pass; you are factorising a matrix whose size depends on the number of variables and constraints in your problem. For a layer inside a training loop that is called on every batch, that matrix is being factorised on every step, and the problem size is the thing that determines whether the approach is viable.

The three frameworks behave the same way. The README says the PyTorch, JAX and MLX layers are functionally equivalent, and the code supports that. The PyTorch version takes tensors and returns tensors, so backward works through the usual autograd machinery. The JAX version uses jax.grad over a lambda that calls the layer, with the PRNG key split to generate the parameter values, so it composes with jit and vmap in the way you would expect. MLX is listed as a third option with functionally equivalent layers, and there is an examples/mlx/ directory, though the README's code examples cover only PyTorch and JAX.

The fast path is Moreau, and you are told to check its licence

The GPU story in this README is not what you would expect from an Apache-2.0 library, and the reason is stated in a sentence that is easy to skim.

CVXPYlayers 1.0 supports GPU acceleration with the Moreau and CuClarabel backends. And then, in the installation section: for the best performance on CPU and GPU, install Moreau via the moreau extra. Before installing, review the Moreau installation guide for license terms and access requirements.

License terms and access requirements. That is a different category of dependency from everything else in this project, and the README does not elaborate because elaborating is not its job. The implication is concrete: the fastest supported path to GPU acceleration is a third-party solver that is not covered by this library's Apache-2.0 licence, and it may require an account, a key or a commercial agreement.

That is not a criticism. It is how solver ecosystems work, and a research library that needs a fast conic solver on a GPU is going to point at whichever one is fast. But an evaluator has to notice that the recommended path is gated, and that installing it is a decision with procurement implications rather than a pip install.

The install itself is an extra, which is good packaging:

bash
pip install cvxpylayers[moreau]

And the CUDA wheel is a separate second step, chosen by CUDA major version:

bash
pip install "moreau[cuda12]"   # or moreau[cuda13]

The pyproject manifest records this as moreau>=0.3.0 in the optional-dependencies table, alongside torch>=2.0.0 for the torch extra, mlx>=0.27.1 for mlx, mlx-metal>=0.27.1 for metal, and jax>=0.4.0 with jaxlib>=0.4.0 for jax. So the frameworks are genuinely optional rather than mandatory, and installing cvxpylayers bare gives you the core with no framework at all until you pick one. There is also an all extra for people who want everything, which notably includes diffqcp>=0.4.4 rather than moreau.

One small thing in the manifest worth noting. The project description reads Solve and differentiate Convex Optimization problems on the GPU, which overstates the default: the GPU is an extra path, not the default one, and the default path is a CPU conic solve plus a CPU sparse factorisation per backward pass. The description is the package metadata that appears on PyPI, so it is the first thing many potential users read.

The homepage field is also worth a glance. The project.urls table sets Homepage to https://github.com/cvxpy/cvxpylayers, which is the repository rather than a documentation site. There is no separate docs domain, and the Documentation entry points at the same URL. So the documentation is the repository's docs/ directory plus the README plus the linked paper and blog post, which is a defensible arrangement for a library with this much paper material but means the README is the front page for anyone arriving from PyPI.

The open-source GPU path is six moving parts and one git checkout

The README offers CuClarabel as an open-source alternative to Moreau, and the list of what you need is the most honest thing in the installation section.

As an open-source alternative, you can use CuClarabel for GPU acceleration. This requires installing Julia and several additional packages: Julia itself, CuClarabel, which is a branch of Clarabel.jl, juliacall, which is the Julia-to-Python bridge from PythonCall, cupy, diffqcp, and lineax from main.

Six items. Take them one at a time and the stack makes sense. CuClarabel is a Julia implementation of an interior point conic solver with CUDA support, and it lives in the Julia ecosystem rather than the Python one. juliacall is how you call Julia from Python without a subprocess, so the solver runs in-process. cupy is the GPU array library, so the data does not make a round trip to the host. diffqcp is the derivative package for this path, and it is a different package from diffcp: the core dependency is diffcp for the conic case, while diffqcp appears in the all extra at version 0.4.4 or newer. And lineax is the automatic differentiation library for the Julia side, needed because the KKT solve has to be differentiated on the Julia side where the solver lives.

And then: lineax from main, with the install command given as an example using uv add "lineax @ git+https://github.com/patrick-kidger/lineax.git". Not a released version. The branch head.

That is the honest state of an emerging integration, and it is worth reading as such. A dependency pinned to a moving branch means your environment is not reproducible, that a lineax commit can break your build without any change on your side, and that you are in the position of reporting bugs against unreleased code. The maintainers presumably track main because the needed functionality is not yet released. The alternative would be not offering the path at all.

The comparison with Moreau is stark and is the practical summary of this section. Moreau is one pip extra plus one CUDA wheel, subject to licence terms and access requirements. CuClarabel is a Julia runtime, a language bridge, a GPU array library, an alternate derivative package, an unreleased AD library, and the fork of a solver. Both get you GPU acceleration. One is a procurement question and the other is an engineering project.

The PyTorch code example for the CuClarabel path is short enough to see the whole surface, and it is worth reading for what it does not say:

python
layer = CvxpyLayer(problem, parameters=[A, b], variables=[x], solver=cp.CUCLARABEL).to(device)
A_tch = torch.randn(m, n, requires_grad=True, device=device)
b_tch = torch.randn(m, requires_grad=True, device=device)

The difference from the CPU example is a solver=cp.CUCLARABEL argument and a .to(device) call, with the tensors created directly on the CUDA device. The API is otherwise identical, which is the point of the design. What is not shown is the Julia side: how long the first call takes while Julia precompiles, whether the Julia solver is actually running on the GPU or on the CPU behind the bridge, and what the failure looks like when juliacall cannot find Julia. None of that is documented, and all of it is what you would discover.

Dual variables and log-log convex programs, the two extensions past the paper

The NeurIPS 2019 paper is the reason this library exists, and two of its features go beyond what the paper covers. Both are worth knowing about because they change what you can build.

The first is dual variables. CVXPYlayers can return constraint dual variables, which are the Lagrange multipliers, alongside the primal solution:

python
eq_con = cp.sum(x) == b
prob = cp.Problem(cp.Minimize(c @ x), [eq_con, x >= 0])

# Request both primal and dual variables
layer = CvxpyLayer(prob, parameters=[c, b], variables=[x, eq_con.dual_variables[0]])

x_star, eq_dual = layer(c_tch, b_tch)

The mechanism is that you name the dual variable alongside the primal variables in the variables list, and the layer returns both. That is a small API change with a large consequence. With duals available and differentiable, you can do sensitivity analysis on a constraint: how much would the optimal value change if I tightened this equality, and what is the marginal price of this constraint in the current solution. That is the kind of question a convex program is uniquely good at answering, and without duals you would have to re-solve the problem for each perturbation.

The important operational detail is that requesting duals changes the shape of the return value. The CPU example destructures a one-tuple, (solution,) = layer(A_tch, b_tch), and the dual example destructures a pair, x_star, eq_dual. So the call site changes, which means adding duals later is not a local edit, and anyone reading your code has to know which form is in play.

The second is log-log convex programs. These are the class of problems that generalise geometric programs, and they are the natural home for problems with sign changes, since a geometric program requires every variable to be positive. Use the gp=True keyword argument when constructing a CvxpyLayer for an LLCP, and the example is a three-variable problem with an objective of 1 over the product of the variables and a constraint that is a product of sums.

The assertion for that path is stricter than the standard one. The DPP example asserts problem.is_dpp(), and the LLCP example asserts problem.is_dgp(dpp=True). Asking for both means the problem has to be a disciplined geometric program and additionally disciplined-parametrized. That is a real narrowing, and it is the right check to run, because log-log convex composition does not automatically preserve the parametrization discipline.

So the library covers three regimes: DPP-compliant convex problems, DPP-compliant log-log convex problems, and both with dual variables. That is a wider surface than the paper, and each of the three has a precondition you must satisfy before you discover it from an exception.

SciPy is a required dependency the README's list omits

The README has a Dependencies section, and it is incomplete in a way that could cost an afternoon.

The list it gives is: Python 3.11 or newer, NumPy 1.22.4 or newer, CVXPY 1.9.0 or newer, and diffcp 1.1.0 or newer. Then it says additionally, install one of PyTorch 2.0 or newer, JAX 0.4.0 or newer, or MLX.

The pyproject manifest lists four runtime dependencies, and the fourth is not in that list:

python
dependencies = [
    "cvxpy>=1.9.0",
    "diffcp>=1.1.0",
    "numpy>=1.22.4",
    "scipy>=1.13.0",
]

SciPy at 1.13.0 or newer. pip install cvxpylayers pulls it in, so this is not a broken install; it is a documentation gap. And it is a gap in the dependency most likely to matter for anyone doing something unusual, because SciPy is what provides the sparse linear algebra the KKT solve needs.

That matters for a specific practical reason. If you are on a platform where SciPy has no wheel and pip tries to build it from source, you will hit a long compile or a failure, and the README gives you no reason to expect SciPy to be involved. The most common places for that are ARM Linux boards, some Raspberry Pi configurations, and musl-based Linux distributions. On those platforms, knowing in advance that a sparse solver is in the dependency set changes the conversation from a build mystery to a known constraint.

The rest of the manifest is worth reading alongside, because it is well organised. The optional-dependencies table separates torch, mlx, metal, jax, moreau and all, so a user picks exactly one framework. The classifiers declare Python 3.11, 3.12 and 3.13, which matches the requires-python floor of 3.11 and means the project is not yet claiming 3.14. And the Development Status classifier reads 5 - Production/Stable, which is a claim worth noting given that the package is at version 1.2.0 and the GPU path depends on a library installed from a moving branch.

The licence field is the one place the manifest uses an older form. It reads license = { text = "Apache-2.0" } rather than the SPDX string license = "Apache-2.0". The table form is still accepted and the classifiers list includes the Apache Software License classifier, so the licence is unambiguous. The SPDX string is the form newer packaging guidance recommends, and the table form is on a deprecation path in some tooling. It is a cosmetic issue for a user and a small piece of housekeeping for a maintainer.

The people are worth noting too. The authors list is A. Agrawal, B. Amos, S. Barratt and S. Diamond, and there is a single maintainer, P. Nobel, with a berkeley.edu address. Boyd is the author of the paper the README links but is not in the package author list, which fits a paper with a large author group and a smaller maintaining group.

The version comes from git tags, and the type checker is Pyright

Two build details that say how seriously this project takes itself, neither of which appears in the README.

The first is the versioning. The pyproject manifest sets dynamic = ['version'] and the build backend is hatchling with hatch-vcs. hatch-vcs derives the version from the repository's own git history, which means the version number in a built artefact is the tag it was built from, and a build from a commit with no tag gets a development version rather than a stale number someone forgot to bump.

That is a different discipline from the manual approach, and it is verifiable in the release history. The three most recent tags are v1.0.3 on 2026-02-27, v1.0.4 on 2026-03-23, and v1.2.0 on 2026-05-19. The patch, the patch and the minor, with the 1.0 line starting in early 2026 and a minor feature release by May. The last push was on 2026-09-24, so the default branch is four months past the newest tag, and because the version is derived rather than typed, a pip install of the default branch gets an unambiguous dev version rather than a stale 1.2.0.

The second is the type checking. The repository has a pyrightconfig.json, so the project is type-checked with Pyright, Microsoft's checker, rather than mypy. It also has a .pre-commit-config.yaml, so the check runs before a commit lands. And CLAUDE.md sits at the root alongside PROCEDURES.md and docs/.

Choosing Pyright over mypy is a small decision with real consequences for a scientific library. Pyright is stricter by default about unannotated returns and implicit Any, and in a codebase that manipulates arrays with shape-sensitive semantics, the type checker is the only thing standing between a user and a silent shape bug. It is also the checker that understands numpy and torch idioms better out of the box.

The repository layout supports the same reading. There is a tests/ directory, a tools/ directory, and an examples/ directory split into jax/, mlx/ and torch/ subdirectories with its own README. So the three frameworks are not just claimed to be equivalent in prose; they have parallel example trees, which is the only way to keep that claim honest.

What is absent is equally telling. There is no LICENSE-adjacent CHANGELOG at the root, no CONTRIBUTING file, and no CI configuration in the top-level listing beyond .github. The procedural documentation is in PROCEDURES.md rather than in a conventional contributing guide. For a library from a research group where the paper is the primary artefact and the code is the accompanying implementation, that allocation makes sense: the paper documents the method, the procedures document how they work on it, and the examples document how to use it. The user-facing cost is that a contributor looking for the usual onboarding files has to find PROCEDURES.md on their own.

Editorial conclusion

Use cvxpylayers if you have a convex optimization problem embedded in a learning pipeline and the problem satisfies disciplined parametrized programming, because the backward pass is a KKT linear solve rather than a finite-difference approximation, and DPP gives a well-defined derivative where an unrolled solver would not have one. Do not adopt it for a problem you have not checked against DPP, since CVXPY accepts a strict superset and the failure is an assertion rather than a diagnostic. Do not plan around the GPU path until you have read the Moreau licence terms and access requirements, because the recommended fast path is not an open-source dependency, and the alternative needs Julia, juliacall, cupy, CuClarabel, diffqcp and a lineax checkout from git main. Verify five things. Assert your problem passes is_dpp() before you build the layer, and is_dgp(dpp=True) if you are using the log-log convex path. Install one framework rather than all of them, since torch, mlx, metal, jax and moreau are optional extras and diffcp is already a core dependency. Note that scipy is a required dependency even though the README's dependency list omits it. Pin diffcp, because it is the component doing the actual differentiation and its behaviour is what determines whether your gradient is correct. And decide early whether you need dual variables, because requesting them changes the output tuple shape and is not something you can add later without touching the call site. The deciding fact is that this is a research library from the paper's authors, at version 1.2.0, with a mathematically correct backward pass and a hardware story that is still being worked out.

Frequently asked questions

How do I install cvxpylayers?

Run pip install cvxpylayers, then install one framework separately, since PyTorch 2.0+, JAX 0.4.0+ and MLX are all optional extras. The core requirements are Python 3.11+, NumPy 1.22.4+, CVXPY 1.9.0+, diffcp 1.1.0+ and SciPy 1.13.0+, though the README's dependency list omits SciPy even though it is in the manifest.

What is DPP and why does cvxpylayers require it?

DPP is disciplined parametrized programming, a CVXPY modelling discipline that restricts how parameters may appear in a problem. The problems you put in a CvxpyLayer must satisfy it, and the README's examples assert problem.is_dpp() before constructing the layer. The set of problems CVXPY can solve is a strict superset of the set that can be differentiated this way.

How does cvxpylayers compute gradients?

The forward pass solves the convex problem and the backward pass computes the derivative of the solution with respect to the parameters. The derivative comes from diffcp, which solves the KKT system of the conic program rather than unrolling the solver's iterations, so the gradient is the exact derivative of the solution map rather than an approximation of the gradient of an approximate iterate. SciPy provides the sparse linear algebra.

How do I get GPU acceleration?

Two paths. The recommended one is the moreau extra, pip install cvxpylayers[moreau] plus a matching CUDA wheel such as pip install "moreau[cuda12]", and the README tells you to review Moreau's installation guide for license terms and access requirements first. The open-source alternative is CuClarabel, which needs Julia, CuClarabel, juliacall, cupy, diffqcp and lineax installed from its git main branch.

Can cvxpylayers return Lagrange multipliers?

Yes. Name the dual variable alongside the primal variables when constructing the layer, for example variables=[x, eq_con.dual_variables[0]], and the call returns both, as in x_star, eq_dual = layer(c_tch, b_tch). Note that requesting duals changes the return arity, so the call site differs from the primal-only form which destructures a one-tuple.

Can cvxpylayers handle geometric or log-log convex problems?

Yes, by passing gp=True when constructing the layer. Log-log convex programs generalise geometric programs and allow sign changes that a geometric program forbids. The precondition is stricter than the standard one: the examples assert problem.is_dgp(dpp=True), meaning the problem must be both a disciplined geometric program and disciplined-parametrized.

Official sources

  1. cvxpy/cvxpylayers on GitHub
  2. Issues
  3. License: Apache-2.0
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/cvxpy-cvxpylayers.svg)](https://hysenlabs.com/projects/cvxpy-cvxpylayers)