soup copies BeautifulSoup's interface, including its global state
Web Scraper in Go, similar to BeautifulSoup
At a glance
- What is it?
- The Go package is small and the readme is essentially the function list, which makes it easy to evaluate. What is worth reading carefully is that the fetching functions are package-level globals, that every navigation method returns an error inside the node rather than returning one, and that the whole library fits in one source file.
- Who is it for?
- Adopt soup if you want a scraping API that reads like the Python library you already know and you are writing a short script rather than a service, since the interface transfers directly and the whole thing is one file you can read in a minute. Do not adopt it for anything concurrent, because the fetch functions and the header and cookie setters are package-level globals, which means two goroutines sharing a client share their configuration.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 61 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The readme is a function signature list, which is unusually honest
Most library readmes open with what the library is for. This one opens with a comparable phrase and then gives you the exported surface, each entry with a one-line description of its arguments and return. That is a design decision about documentation, and it has an obvious benefit: you can see the whole API without scrolling, and you can see what is missing. Counting the fetch functions, there are four: one that takes a URL, one that takes a URL and your own HTTP client, one that posts a payload with a declared body type, and one that posts form-encoded values. Counting the navigation functions, there are four finders and four sibling walkers, plus a children accessor, an attributes accessor, and two text accessors with different behaviour. Then a debug switch and a serialiser. That is about twenty functions, and the count is the point: this is a surface small enough to hold in your head. For a library whose description is similar to a well-known Python package, the absence of anything else is the message. What you get is traversal and extraction over a parsed document, plus the fetching to get there. What you do not get is a selector engine, a session abstraction, a retry policy, an encoding detector, or a way to run two configurations at once. Every one of those omissions is a decision, and they were made to keep the interface mirrorable from another language.
Package-level globals for headers and cookies, which is the design's real cost
Two entries in the signature list are variables rather than functions, and they are the ones to think about first. One is a map of headers, described as an alternative to calling a setter individually. The other is a map of cookies, described the same way. Alongside them are two functions that set a single header and a single cookie, with comments saying they apply to the request made by the fetch function. So the configuration for an HTTP request lives in package-level state, mutated by exported functions, and read by the fetch call. In a single-threaded script that is convenient and costs nothing. In a program with concurrency it is a bug waiting to happen, because two goroutines that want different authentication headers will overwrite each other's values between the set and the send. There is also no per-request configuration on the fetch functions themselves, except for the variant that takes your own client, and that is the escape hatch: build a client with a transport that injects the headers you want and use the client-taking variant. The readme does not suggest this, and it does not warn about it, so a user writing concurrent code will find it out the hard way. If you need per-request headers, this is the pattern to use from the start.
Errors live in the node, and a node is a struct with three fields
Installing is a single module fetch, and the readme's example is a full program rather than a fragment:
package main
import (
"fmt"
"os"
"github.com/anaskhan96/soup"
)
func main() {
resp, err := soup.Get("https://xkcd.com")
if err != nil {
os.Exit(1)
}
doc := soup.HTMLParse(resp)
links := doc.Find("div", "id", "comicLinks").FindAll("a")
for _, link := range links {
fmt.Println(link.Text(), "| Link :", link.Attrs()["href"])
}
}The core type is described as a struct with three fields, and one of them is an error. There is a pointer to the current node, a value field holding the tag name for an element node or the text for a text node, and an error field that is nil when nothing went wrong. Alongside it is a field whose type is an enumeration of what can go wrong, and the readme enumerates the cases: unable to parse, element not found, no next sibling, no previous sibling, no next element sibling, no previous element sibling, unable to create a get request, error in a get request, and error reading the response. That is a well-chosen set, and the design of carrying the error on the node is what makes the fluent interface possible. A chain like find a container then find all links inside it reads as one expression, and the error from the first step travels with the node into the second. The cost is that a chain which fails partway returns something, and you have to check. The readme's example does not check anything beyond the initial fetch error, which is a fair warning about how the API is meant to be used and also an example of how it fails silently. An experienced user will check the error field after each step; a user copying the example will not.
Text versus full text, and the two text accessors that look identical
There are two accessors that return a string and the descriptions are almost the same. One is described as returning the full text inside a non-nested tag, and the first half in a nested one. The other returns the full text inside a nested or non-nested tag. So the difference is that the first one truncates when the element contains children, and the second does not. That is a real distinction and it is the kind of thing that produces a bug when you switch from one to the other expecting a rename, so it is worth knowing before you need it. The attributes accessor returns a map of the element's attributes, and the readme's example uses it directly to pull a link target out of a found anchor, which is the shape of most scraping code. The serialiser returns the HTML for the specific element, which is what you want for extracting a fragment rather than a value. And the children accessor returns direct children only, not descendants, which is the correct semantics for a DOM API and the thing that distinguishes it from a flat find-all. Taken together the extraction surface is: move through the tree, filter by tag and attribute, read text, read attributes, read children, serialise. It is enough for a scraper and it is deliberately not more than that.
One source file, two dependencies, and a major version that appeared this year
The repository is strikingly small and that is the third thing worth reading. At the top level there is one source file, one test file, a module file, a lock file, a changelog, a licence file, a readme, and an examples directory. The module declares a single direct dependency beyond a test assertion library, and that dependency is the HTML parsing and network library from the standard extended library, at a specific version. The indirect dependencies in the module file are exactly what you would expect from that one: a structured data differ used by the assertion library, a difference library, an internationalisation library, and a YAML library. So the dependency surface is one parsing library and its transitive requirements, which for a scraper is about as small as it can be while still parsing HTML. The examples directory has four subdirectories, and their names are informative: one for a non-English page, one for error handling, one for a live data feed, and one for extracting from a specific site. The last one is what the readme's example does, and the third is the interesting one, because a scraper that only handles static pages is a different thing from one that handles a feed. The release history is three versions in the same major, with the newest two a major bump and a patch within hours of each other in July, followed by a patch in August. A second major in a project with this much history suggests an interface change, and the readme has no migration notes.
Editorial conclusion
Adopt soup if you want a scraping API that reads like the Python library you already know and you are writing a short script rather than a service, since the interface transfers directly and the whole thing is one file you can read in a minute. Do not adopt it for anything concurrent, because the fetch functions and the header and cookie setters are package-level globals, which means two goroutines sharing a client share their configuration. Four things to verify. That your tasks are single-threaded, because the readme's own signature list shows the headers and cookies as package variables set by mutating functions, and that is the design. How you handle errors, because they are carried inside the node rather than returned, so a failed lookup produces a node whose error field is set and whose other methods are probably not worth calling. Whether the parser handles what you are scraping, because the dependency list has one HTML parser and the readme says nothing about malformed markup, encoding, or how it treats a document the parser had to repair. And that the interface has not changed under you, since the release history shows a major version bump and the readme does not include migration notes. The licence is MIT, version 2.0.2 was released on 2026-08-01, and the last push was the same day.
Frequently asked questions
What does the soup Go package do?
It fetches a page over HTTP with a get or a post, parses the returned markup into a node tree, and gives you an interface similar to BeautifulSoup: find one element or all elements by tag and attribute, strict variants for exact attribute matching, sibling traversal in four directions, direct children, attributes, two kinds of text extraction, and serialisation of an element back to markup.
How do I install and use soup?
Fetch the module with the standard Go package command, then call the fetch function with a URL, parse the returned string, and chain find calls. The readme's example fetches a page, parses it, finds a container by attribute, finds all links inside it, and prints each link's text and its target from the attribute map.
How are errors reported in soup?
Errors are carried inside the node rather than returned. The node struct has a pointer, a value field holding the tag or text, and an error field that is nil on success, with a type field drawn from an enumerated set covering parse failure, element not found, the four missing-sibling cases, and three request-level failures.
Can I set headers and cookies per request in soup?
Not directly. The readme lists package-level variables for headers and cookies plus functions that set a single pair, and the comments say they apply to the request made by the fetch function. The variant that accepts your own HTTP client is the way to get per-request behaviour, and the readme does not discuss the concurrency implications of the package-level state.
What licence is soup released under?
MIT, with the licence file at the repository root. Version 2.0.2 was released on 2026-08-01, the same day as the last push, preceded by a second major version and a patch in July 2026. The implementation is a single source file, and the module has one direct dependency beyond a test library.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/anaskhan96-soup)