# PyGAD: a genetic algorithm library for Python that plugs into Keras and PyTorch

> PyGAD wraps selection, crossover and mutation behind one pygad.GA object and lets you supply only a fitness function. It is a good fit for small search spaces and for evolving neural network weights; it is not a replacement for gradient descent or for a constrained MIP solver.

**ahmedfgad/GeneticAlgorithmPython** — Source code of PyGAD, a Python 3 library for building the genetic algorithm and training machine learning algorithms (Keras & PyTorch).

- Repository: https://github.com/ahmedfgad/GeneticAlgorithmPython
- Website: https://pygad.readthedocs.io
- Stars: 2,226 · Forks: 500
- Language: Python
- License: BSD-3-Clause
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/ahmedfgad-geneticalgorithmpython

## What PyGAD actually replaces in your code

Writing a genetic algorithm by hand is mostly bookkeeping. You need a population representation, a parent selection rule, a crossover operator, a mutation operator, a way to keep the best individual, and a loop that repeats all of it. PyGAD supplies those pieces and asks you for one thing: a fitness function that takes a candidate solution and returns a number. The README states that PyGAD supports different types of crossover, mutation and parent selection, and that the library is aimed at both single-objective and multi-objective problems.

The audience is narrow but real. If you are optimizing a function with a handful to a few hundred parameters, if you are tuning a small neural network's weights, or if you are solving a combinatorial puzzle such as a travelling salesman instance, PyGAD gives you a working loop in a few lines. The repository ships an examples/ directory with entries such as example.py, example_multi_objective.py, example_multi_objective_nsga3.py, example_parallel_processing.py and example_gene_constraint.py, plus a notebook for the travelling salesman problem. Those filenames tell you the intended scope better than any marketing line: small, self-contained optimization experiments.

Where it is misapplied is when people reach for it because they cannot differentiate their objective. A genetic algorithm spends thousands of evaluations to find what a gradient method finds in dozens, so on a differentiable objective with many parameters the trade is bad. PyGAD does not stop you from making that mistake.

## The pygad.GA lifecycle and the callbacks that expose it

A pygad.GA instance runs a fixed sequence of stages, and the README publishes a lifecycle diagram plus a callback for each stage. The callbacks are on_start, on_fitness, on_parents, on_crossover, on_mutation, on_generation and on_stop. The README states that PyGAD stops when all generations are completed or when the function passed to on_generation returns the string stop. That is the only early-termination hook the README documents.

This design is the library's main architectural decision: instead of subclassing, you pass functions. It keeps the core small and makes the run inspectable. The cost is that state lives on the ga_instance object rather than in your own class, and the callback signatures are fixed by the library, so anything you want to record has to be captured from the arguments you are handed.

The fitness function signature in the README is fitness_func(ga_instance, solution, solution_idx), returning a scalar. Note the first argument. Older PyGAD code in the wild uses fitness_func(solution, solution_idx), and the README does not document whether the two-argument form still works. If you are copying an example from a blog post, check which signature it uses before you debug anything else.

The README example computes fitness as 1.0 / (numpy.abs(output - desired_output) + 0.000001). That epsilon matters. Without it, an exact match divides by zero. PyGAD maximizes fitness, so a minimization problem has to be inverted, and the inversion has to stay positive.

## Installing PyGAD and running a first optimization

PyGAD is on PyPI. The README gives a single install command, and the core install pulls in only numpy and cloudpickle, which is a deliberate choice noted in the README itself.

```bash
pip install pygad
```

Plotting and deep learning are optional extras. The README shows these two forms, and the repository's setup.py declares the same extras under extras_require.

```bash
pip install pygad[visualize]
pip install pygad[deep_learning]
```

The visualize extra adds matplotlib for plot_fitness() and plot_genes(). The deep_learning extra adds keras, tensorflow and torch for pygad.kerasga and pygad.torchga. If you only need the optimizer, skip both; you avoid a TensorFlow install for nothing.

Here is the smallest useful run, adapted from the README example. It searches for a vector whose weighted sum is close to 44.

```python
import pygad
import numpy

function_inputs = [4, -2, 3.5, 5, -11, -4.7]
desired_output = 44

def fitness_func(ga_instance, solution, solution_idx):
    output = numpy.sum(solution * function_inputs)
    return 1.0 / (numpy.abs(output - desired_output) + 0.000001)

ga_instance = pygad.GA(num_generations=3,
                       num_parents_mating=5,
                       fitness_func=fitness_func,
                       sol_per_pop=10,
                       num_genes=len(function_inputs))

ga_instance.run()
```

After run() returns, the instance holds the result. The README's callback example prints stage names; for the answer itself you read the best solution and its fitness off the instance. num_genes must equal the length of the solution vector, and sol_per_pop must be at least num_parents_mating, or the parent selection step has nothing to choose from.

## Where PyGAD gets in the way

The fitness function is called once per individual per generation, in Python. There is no compiled inner loop and no vectorized batch evaluation in the core described by the README. With sol_per_pop=10 and num_generations=3 that is 30 calls and irrelevant; with a population of thousands over thousands of generations, and a fitness function that trains or simulates something, the wall-clock cost is the whole project. The README points to an example_parallel_processing.py file, so parallelism exists, but the README does not describe its overhead or its scaling, and process startup on a fast fitness function can cost more than it saves.

Constraint handling is the second rough edge. The repository has example_gene_constraint.py and example_gene_space.py, so there is machinery for restricting genes, but the README does not explain how a constraint is enforced during crossover or mutation. If your problem has hard feasibility requirements, you should read those two examples before assuming the library will keep your population legal.

Third, the README and the released package disagree. The README's lifecycle example uses the three-argument fitness_func(ga_instance, solution, solution_idx) signature, while the published documentation site is the authoritative reference for the installed version. Treat the README as a tour, not as an API contract.

Finally, PyGAD is the wrong tool when your objective is differentiable and high-dimensional, when you need a certificate of optimality, or when the feasible region is defined by linear constraints that a solver could handle exactly. A genetic algorithm gives you a good-enough answer with no guarantee, and that is the deal.

## PyGAD against DEAP and plain scipy.optimize

DEAP is the other well-known Python evolutionary computation library, and the difference is philosophical. DEAP gives you a toolbox and a creator: you define your individual type, register your own crossover, mutation and selection operators, and assemble a loop. You get full control over representation, including trees and arbitrary containers, at the price of writing the plumbing yourself. PyGAD inverts that. It fixes the representation to a numeric array of genes and gives you a ready pygad.GA object, with customization available through parameters and callbacks rather than through subclassing. If your problem fits a fixed-length numeric vector, PyGAD is faster to a working run. If it does not, DEAP will not fight you.

scipy.optimize is a different comparison entirely. Methods such as differential_evolution also avoid gradients, but they assume a smooth-ish continuous objective and return a result object with a success flag. PyGAD's advantage is the callback lifecycle and the built-in bridge to Keras and PyTorch through pygad.kerasga and pygad.torchga, which the README calls out explicitly. Its disadvantage is that it is a general framework, so you pay framework overhead for problems scipy already solves in one call.

The honest summary: reach for PyGAD when the solution is naturally an array of numbers and you want evolutionary training of a model. Reach for DEAP when the representation is unusual. Reach for scipy when the problem is smooth.

## Maintenance, licence and the cost of upgrading

The repository is not archived. The last push was on 2026-07-09, and the most recent release listed is 3.7.0 from 2026-06-05, preceded by 3.6.0 on 2026-04-08 and 3.5.0 on 2025-07-09. The README describes the library as under active development with features added regularly, and the release cadence through 2026 is consistent with that.

Upgrade cost is the practical question. The project uses semantic versioning in its tags, and the 3.x line has moved three times in roughly fourteen months. The README does not document a deprecation policy or a rollback procedure, so before bumping a minor version in a long-running job you should check the release notes for that tag rather than assume compatibility. Pin the version in your own requirements file.

Licensing is BSD-3-Clause, declared in the LICENSE file, in the badge at the top of the README, and in the REUSE.toml and LICENSES/ entries in the repository root. That is a permissive licence: it allows commercial use and modification, and it requires that the copyright notice and disclaimer be retained in redistributions. Note that pyproject.toml currently carries the license as a file reference with the SPDX string commented out, so tooling that reads the SPDX field may not detect the licence automatically. That is a packaging detail, not a legal one. This is a description of the licence text, not legal advice; read the LICENSE file yourself if the distinction matters to your organization.

## Conclusion

Adopt PyGAD if your problem is a small or medium search space where you can score a candidate cheaply, or if you specifically want evolutionary training of a Keras or PyTorch model. Do not adopt it as a drop-in replacement for gradient descent on a large network, and do not expect it to enforce hard constraints for you. Before committing, verify that your fitness function is deterministic and fast enough for num_generations multiplied by sol_per_pop evaluations, and check the pygad.readthedocs.io page for the version you pin, because the README example and the released API have drifted apart.

## FAQ

### What is a genetic algorithm?

It is an optimization method that keeps a population of candidate solutions, scores each one with a fitness function, then produces the next generation by selecting parents and applying crossover and mutation. PyGAD implements that loop and exposes each stage through callbacks such as on_parents and on_mutation.

### How can I write a genetic algorithm in Python with PyGAD?

Install the library with pip install pygad, write a fitness function that accepts a solution and returns a number, then construct pygad.GA with num_generations, num_parents_mating, fitness_func, sol_per_pop and num_genes, and call run(). The README's lifecycle example is the shortest complete version.

### Is a genetic algorithm AI?

PyGAD is presented as a library for building the genetic algorithm and optimizing machine learning algorithms, and it ships pygad.kerasga and pygad.torchga modules for training Keras and PyTorch models. Whether that counts as AI depends on your definition; the library itself makes no claim beyond optimization.

### How do I write a genetic algorithm step by step?

The README publishes a lifecycle diagram for a pygad.GA instance and a callback for each stage: on_start, on_fitness, on_parents, on_crossover, on_mutation, on_generation and on_stop. Running that callback example prints each stage name in order, which is the clearest way to see the sequence.

### What is the Python genetic algorithm library for optimization?

PyGAD is a Python 3 library for building the genetic algorithm and optimizing machine learning algorithms, distributed on PyPI as pygad. Its core install depends only on numpy and cloudpickle, with matplotlib, Keras and PyTorch available as optional extras.

## Sources

- [ahmedfgad/GeneticAlgorithmPython on GitHub](https://github.com/ahmedfgad/GeneticAlgorithmPython)
- [License: BSD-3-Clause](https://github.com/ahmedfgad/GeneticAlgorithmPython/blob/master/LICENSE)
- [Project website](https://pygad.readthedocs.io)
- [README](https://github.com/ahmedfgad/GeneticAlgorithmPython/blob/master/README.md)
- [Releases](https://github.com/ahmedfgad/GeneticAlgorithmPython/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/ahmedfgad-geneticalgorithmpython
