Model or dataset
spcl/graph-of-thoughts avatar
spcl/graph-of-thoughts

Graph of Thoughts: modelling LLM reasoning as a Graph of Operations

Official Implementation of "Graph of Thoughts: Solving Elaborate Problems with Large Language Models"

2,842 stars218 forksPythonNOASSERTION

At a glance

What is it?
spcl/graph-of-thoughts is the official Python implementation of the GoT paper. It lets you express prompting strategies as an explicit Graph of Operations and execute them against an LLM through a Controller.
Who is it for?
Adopt it if you are researching prompting structures or need to reproduce the paper's experiments, because the Graph of Operations abstraction makes CoT, ToT and GoT variants directly comparable. Do not adopt it as a production orchestration layer: the package version is 0.0.3, the last release tag is v0.0.2 from 2023-09-26, and the only documented LLM configuration path is the Controller README.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Activity is slowing. The repository last received commits 6 months ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The problem GoT solves: prompting strategies that are not chains

Most prompting code is a sequence. You build a prompt, call the model, take the string back, and feed it into the next prompt. That shape is fine for translation or summarisation, where one pass over the input is enough. It breaks down when the task needs several candidate answers that get scored, merged, or refined before a final answer is produced.

Graph of Thoughts treats those structures explicitly. Instead of hard-coding a loop that generates five candidates and picks the best, you declare a Graph of Operations (GoO): Generate, Score, Aggregate, Improve, GroundTruth, and the branching between them. The README describes the framework as solving complex problems by modelling them as a GoO, which is automatically executed with an LLM as the engine. The intended audience is researchers and engineers who want to compare reasoning structures rather than hand-tune one prompt. The README also states that you can implement GoOs resembling previous approaches like CoT or ToT, so the same code path covers the baselines.

How the Graph of Operations, Prompter, Parser and Controller fit together

Four pieces carry the data flow. The GraphOfOperations holds the operations and their edges. A Prompter turns the current thought state into text for the model. A Parser turns the model's text back into a structured thought. The Controller drives the loop: it walks the graph, asks the Prompter for a prompt, calls the language model, hands the response to the Parser, and stores the resulting thought in the graph.

The thought state itself is a dictionary. In the README's sorting example it contains keys such as original, current, method, and phase. That dictionary is what the Prompter reads, so adding a new field to the state is how you thread extra context through the graph without changing the Controller.

The Score operation is where the graph earns its keep. It takes a scoring_function, and the sorting example passes utils.num_errors, so intermediate thoughts are ranked before the next operation runs. GroundTruth then checks the final thought against utils.test_sorting. Because scoring happens inside the graph, you can branch, keep the best k thoughts, and merge them, which is the structural difference from a single chain.

After ctrl.run(), the Controller can serialise the whole graph with ctrl.output_graph("output_cot.json"). That file is the audit trail: it records the thoughts and their scores, not just the final string.

Installing graph_of_thoughts and running the sorting example

The README requires Python 3.8 or newer and says to activate your Python environment first. If you only want to use the library, install it from PyPI:

bash
pip install graph_of_thoughts

If you intend to modify the code, install it in editable mode from source instead:

bash
git clone https://github.com/spcl/graph-of-thoughts.git
cd graph-of-thoughts
pip install -e .

Either path pulls the dependencies declared in pyproject.toml, which include openai, transformers, torch, accelerate, bitsandbytes, matplotlib, numpy, pandas, sympy and scipy. That is a heavy install for a prompting library, and it is the first thing to notice if you only plan to call a hosted API.

Before any code runs, you need an LLM. The README does not describe the config format inline; it points to the Controller README under graph_of_thoughts/controller/ for instructions on configuring the LLM of your choice. The quick-start snippets assume a config.json in the current directory holding an OpenAI API key.

With that file in place, the CoT-shaped baseline is short:

python
from examples.sorting.sorting_032 import SortingPrompter, SortingParser, utils
from graph_of_thoughts import controller, language_models, operations

gop = operations.GraphOfOperations()
gop.append_operation(operations.Generate())
gop.append_operation(operations.Score(scoring_function=utils.num_errors))
gop.append_operation(operations.GroundTruth(utils.test_sorting))

lm = language_models.ChatGPT("config.json", model_name="chatgpt")
ctrl = controller.Controller(lm, gop, SortingPrompter(), SortingParser(),
  {"original": to_be_sorted, "current": "", "method": "cot"})
ctrl.run()
ctrl.output_graph("output_cot.json")

The GoT variant swaps the hand-built graph for got() imported from the same example module and sets method to "got" and phase to 0 in the initial state. The README says you can compare the two runs by inspecting output_cot.json and output_got.json, where the final thought states' scores indicate the number of errors in the sorted list. There is also a module-level entry point if you prefer running examples from the repository root:

bash
python -m examples.sorting.sorting_032
python -m examples.keyword_counting.keyword_counting

The README notes that results are stored in the respective examples sub-directory.

Where GoT is the wrong tool, and what the repository does not tell you

The dependency list is the clearest limitation. torch, transformers, accelerate and bitsandbytes are declared as install requirements even though the quick-start path uses language_models.ChatGPT. Installing the package to call a hosted model therefore drags in a local inference stack. If your environment cannot afford that footprint, the PyPI route is a poor fit.

Cost and latency scale with graph shape, and the README gives no budget guidance. A Generate operation that branches into many thoughts multiplies model calls, and Score adds calls of its own. Nothing in the README documents a cap on branching factor, a retry policy beyond the backoff dependency, or a way to estimate spend before a run.

The package version in pyproject.toml is 0.0.3 while the most recent tagged release listed is v0.0.2 from 2023-09-26, so the version you install from PyPI may not correspond to a release tag. The last push to the repository was on 2026-03-24. Treat the API as unstable: nothing in the README promises backward compatibility between minor versions.

The material also does not document multi-user serving, persistence of graphs to a database, or any HTTP interface. This is a library for constructing and executing a GoO in a Python process, not a service. If you need a hosted endpoint with quotas and observability, you are building that layer yourself.

GoT versus Tree of Thoughts and other reasoning structures

Tree of Thoughts is the closest comparison, and the README names it directly as something you can reimplement inside this framework. The structural difference is the one the paper's title implies: a tree allows one parent per node, so two branches can never be merged back together. A graph allows several thoughts to feed a single downstream operation, which is what makes aggregation and refinement of multiple candidates expressible as a first-class step rather than glue code.

That matters for tasks like merging partial documents or combining sorted sublists, which is why the examples directory ships doc_merge, set_intersection and sorting alongside keyword_counting. If your task is a single chain of dependent edits, a tree or even a plain loop is simpler and cheaper, and the graph machinery buys you nothing.

Against a general orchestration framework, the difference is scope rather than features. GoT does not aim to schedule heterogeneous tools or manage long-running workflows. It gives you one abstraction, the Graph of Operations, and one execution model, the Controller loop, and it keeps the Prompter and Parser pluggable so the prompt format stays under your control.

Maintenance, upgrade cost and the licence question

The repository is not archived, and the last push was on 2026-03-24. There are two tagged releases, v0.0.1 from 2023-08-23 and v0.0.2 from 2023-09-26, while pyproject.toml declares version 0.0.3. Pinning to a git commit is more predictable than pinning to a version number here.

Upgrade cost is dominated by the dependency ranges. The openai bound is >=1.0.0,<2.0.0, so an OpenAI client major release will require a code change in language_models. torch is bounded below 3.0.0 and transformers below 5.0.0, which gives reasonable headroom but means a Python environment shared with other ML tooling can conflict. Because the package is installed as graph_of_thoughts and imported as graph_of_thoughts, there is no namespace collision risk with the many other projects using similar names.

On licensing, the repository metadata reports NOASSERTION, and pyproject.toml points the licence field at the LICENSE file rather than naming an SPDX identifier. The README does not discuss licence terms or commercial use. Read the LICENSE file in the repository root before you depend on it, and treat that file, not the package metadata, as the authoritative statement.

Editorial conclusion

Adopt it if you are researching prompting structures or need to reproduce the paper's experiments, because the Graph of Operations abstraction makes CoT, ToT and GoT variants directly comparable. Do not adopt it as a production orchestration layer: the package version is 0.0.3, the last release tag is v0.0.2 from 2023-09-26, and the only documented LLM configuration path is the Controller README. Before committing, check that examples/sorting/sorting_032 runs end to end with your own config.json, and confirm the LICENSE file terms yourself, since the repository metadata reports NOASSERTION rather than a recognised SPDX identifier.

Frequently asked questions

What is Graph of Thoughts?

It is the official implementation of the paper Graph of Thoughts: Solving Elaborate Problems with Large Language Models. The README describes it as a framework that models a problem as a Graph of Operations and executes that graph automatically with an LLM as the engine.

Can you give me an example of a knowledge graph?

The repository does not ship a knowledge graph example. Its examples directory covers sorting, keyword counting, set intersection and document merging, and the README does not describe a knowledge graph use case.

What is a graph example in Graph of Thoughts?

The README's sorting example is the concrete one: a GraphOfOperations with Generate, Score using utils.num_errors, and GroundTruth using utils.test_sorting, run through a Controller and written to output_got.json.

Can AI explain its reasoning?

The framework records intermediate thoughts rather than only a final answer. Running ctrl.output_graph writes the graph to a JSON file such as output_cot.json, and the README says the final thought states' scores indicate the number of errors in the sorted list.

Official sources

  1. Issues
  2. Project website
  3. README
  4. Releases
  5. spcl/graph-of-thoughts on GitHub
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/spcl-graph-of-thoughts.svg)](https://hysenlabs.com/projects/spcl-graph-of-thoughts)