Model or dataset
MODSetter/SurfSense avatar
MODSetter/SurfSense

SurfSense: the open-source NotebookLM alternative that moved on from NotebookLM

Open-source NotebookLM alternative. Research the open web with live data(Reddit, YT, IG, TikTok, Indeed, Google Search, Maps etc) through one platform, API or MCP server. Join our Discord: https://discord.gg/ejRNvftDp9

16,295 stars1,568 forksPythonNOASSERTION

At a glance

What is it?
SurfSense started as a self-hosted knowledge base with cited chat and now sells itself mainly as a live-data layer for agents. The connectors are real, the pivot is documented, and the licence file is the first thing to check.
Who is it for?
Adopt SurfSense if your agents need structured live web data and you want the option to self-host the whole stack. Do not adopt it if you want a drop-in NotebookLM replacement with a stable scope, since the README says the project has pointed its energy at open web research and the older knowledge-base features are kept rather than developed.
Can I use it commercially?
Check first. The repository uses a licence we do not classify automatically, so read its LICENSE file before any commercial use.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.

DEEP OPEN-SOURCE ANALYSIS

What SurfSense is for, and who is actually meant to run it

Ask a capable agent what Reddit has said about a product since launch and it has nowhere reliable to look. That is the problem SurfSense names in its own README. Official platform APIs are described there as rate-limited, priced for enterprises, or missing entirely; scraping plumbing is called brittle; and driving a browser with an LLM is said to burn minutes and tokens per page. SurfSense answers with a catalog of connectors that return structured JSON instead: posts, comments, transcripts, reviews, SERPs, pages.

The audience has shifted. The project was built as a general research agent over your own knowledge, and the README is explicit that this chapter has closed: reasoning over a static index is now something capable agents do out of the box. The stated target is agents that need live data and the workflow around it. If you are shopping for a self-hosted place to upload PDFs and chat with citations, the features still exist, but the project's attention is elsewhere. If you are building an agent that has to answer questions about what real people said on Reddit, TikTok, or Google Maps this week, you are the intended user.

One typed surface, an MCP server, and an agent harness

The mechanism is deliberately unglamorous. Every connector is a REST endpoint that returns structured JSON, so a Reddit scrape returns posts and comments rather than an HTML blob you have to parse. The README frames this as removing three failure modes at once: rate-limit roulette, HTML parsing, and the browser loop.

On top of that sits an MCP server. The README says it exposes every connector as a native tool with names like surfsense_reddit_scrape and surfsense_google_search, so Claude, Cursor, or another agent framework can call them directly. The third layer is what the README calls an agent harness: retries, structured output, and credit metering are built in, which means the retry policy and the billing counter live on the server rather than in your agent loop.

Billing follows the same per-unit logic. Connectors bill per item actually returned, crawls bill per page successfully fetched, and the README states failed calls are never billed. Self-hosted installs run with billing off. That last sentence is the one that matters most for anyone evaluating this as infrastructure rather than as a hosted API.

Installing SurfSense and making a first Reddit call

The README's quick start does not walk through a full self-hosted install; it goes straight to calling a connector, and it points at the connector pages for copy-paste examples in Python, JavaScript, Go, PHP, Ruby, Java, and C#. The repository layout tells you where the pieces live: surfsense_backend, surfsense_web, surfsense_mcp, surfsense_desktop, surfsense_browser_extension, surfsense_obsidian, and a docker/ directory. The package.json at the root is private and pins [email protected] as the package manager, so the web side expects pnpm rather than npm.

The README's Reddit example is a single POST. It references three variables by name, SURFSENSE_API_URL, SURFSENSE_API_KEY, and WORKSPACE_ID, and the payload below is the one the README gives, including the community, sort, and time_filter keys.

bash
curl -X POST "$SURFSENSE_API_URL/workspaces/$WORKSPACE_ID/scrapers/reddit/scrape" \
  -H "Authorization: Bearer $SURFSENSE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "search_queries": ["your brand"],
    "community": "SaaS",
    "sort": "top",
    "time_filter": "week"
  }'

What you should see is structured JSON covering posts and comments for the query, not a rendered page. If you would rather hand the same capability to an agent, the README's MCP configuration is a single server entry pointing at the hosted MCP endpoint, with the API key passed as a bearer token in the headers.

json
{
  "mcpServers": {
    "surfsense": {
      "url": "https://mcp.surfsense.com/mcp",
      "headers": { "Authorization": "Bearer ${SURFSENSE_API_KEY}" }
    }
  }
}

Add that block to your client's MCP configuration and the connector tools appear alongside the client's own tools. For self-hosting, the README does not spell out the docker compose steps in the section shown here, so treat the docker/ directory and the docs/ directory as the sources to read before you commit to a deployment.

Where SurfSense gets in the way

The first limitation is scope drift, and the project is honest about it. The README's own note says the team spent a couple of months building SurfSense as a general research agent and has now redirected its energy. Existing features keep working, but a tool whose center of gravity has moved is a different bet than one that is being deepened in place.

The second is the licence. The repository's licence field reads NOASSERTION, which means the licence could not be identified automatically from the repository metadata. The README says self-hosting stays free and open source, but that sentence is not a licence grant. Before you build a product on top of the backend, open the LICENSE file and read it, and if your legal position depends on a specific OSI-approved licence, confirm it rather than assuming it from the marketing line.

The third is the connector model itself. SurfSense solves rate limits by not using the official APIs, which means the data you get is whatever the connector can extract from public pages. That is a reasonable trade for social listening, and a bad one if your use case depends on a platform's official data guarantees, its terms for programmatic access, or fields the public page does not expose. If you need contractual data provenance, this is the wrong tool. The README also does not document rollback or version pinning for self-hosted upgrades in the section shown, which matters given the release cadence.

SurfSense against NotebookLM and Open Notebook

The comparison the project invites is with NotebookLM, and the honest answer is that they now do different jobs. NotebookLM is a hosted product for reasoning over sources you supply. SurfSense, as described in its README, is a platform for pulling sources in from places NotebookLM does not reach: Reddit, YouTube, Instagram, TikTok, Amazon, Walmart, Google Maps, Google Search, and Indeed, plus a web crawl connector and external MCP connectors for Notion, Slack, and Jira.

The difference in approach is where the data comes from and who drives. NotebookLM waits for you to upload. SurfSense goes and fetches, on demand or on a schedule, and returns typed JSON that an agent can consume without a human in the loop. The README also mentions scheduled and event-triggered agents that turn findings into briefs and alerts, which is a workflow NotebookLM does not present.

Against Open Notebook, the axis is the same: self-hosted notebook tools generally assume you bring the corpus. SurfSense's self-hostable install keeps that model but adds the connector layer and the MCP surface. If your problem is organizing documents you already have, a notebook tool is the simpler answer. If your problem is getting data out of platforms that do not want to give it to you, the connector layer is the whole point.

Release cadence, upgrade cost, and the licence question

The last push to the repository was on 2026-09-10, and the most recent tagged release is v0.0.39 from 2026-08-29, preceded by v0.0.38 on 2026-08-21 and v0.0.37 on 2026-08-20. That is a fast, pre-1.0 cadence: three releases in roughly nine days in August, with the version number still in the 0.0.x range. The practical consequence is that self-hosting means tracking a moving target. The repository carries a VERSION file and a versions.json alongside the code, which suggests version bookkeeping is part of the project's own process.

For upgrade cost, the honest position is that the README does not document a rollback procedure or a supported upgrade path for self-hosted installs. If you deploy this, pin a specific release tag rather than tracking main, and read the changelog before moving. The README links to the changelog at surfsense.com/changelog for the announcement about the pivot, so that page is the natural place to check what changed between tags.

On licence, the repository is flagged NOASSERTION, and the README states self-hosting stays free and open source. Those two statements are not the same thing. Read LICENSE before you redistribute anything or build a commercial product on the backend. This is not legal advice, and the file is short enough that reading it takes less time than asking.

Editorial conclusion

Adopt SurfSense if your agents need structured live web data and you want the option to self-host the whole stack. Do not adopt it if you want a drop-in NotebookLM replacement with a stable scope, since the README says the project has pointed its energy at open web research and the older knowledge-base features are kept rather than developed. Before installing, read LICENSE, because the repository is flagged NOASSERTION, and check the docker/ directory against your own deployment requirements.

Frequently asked questions

What is SurfSense AI?

SurfSense is an open-source research platform that connects AI agents to live web data. Its README describes it as a NotebookLM alternative with connectors for Reddit, YouTube, Instagram, TikTok, Amazon, Walmart, Google Maps, Google Search, Indeed, and any page on the open web, exposed through a REST API and an MCP server.

How do I install SurfSense?

The README's quick start does not give a full self-hosted install walkthrough, but the repository contains a docker/ directory and a docs/ directory, and the root package.json pins [email protected] for the web side. The README's own example starts from an already-running instance and calls a connector with curl.

How do I use SurfSense?

You either POST to a connector endpoint with your API key, workspace ID, and a JSON body, or you add the SurfSense MCP server to an agent client such as Claude or Cursor and let the agent call tools like surfsense_reddit_scrape. The README gives both the curl call and the MCP configuration block.

Is SurfSense free?

The README states that self-hosted installs run with billing off, and that hosted use is pay as you go, billed per item returned by connectors and per page successfully fetched by crawls, with failed calls never billed. It also says self-hosting stays free and open source.

Is SurfSense safe?

The README does not make security claims, so there is nothing to verify against. What it does say is that the project is open source and self-hostable, which lets you keep research on your own infrastructure. Note that the repository's licence field reads NOASSERTION, so read the LICENSE file before deploying.

How does SurfSense compare with NotebookLM?

NotebookLM reasons over sources you supply; SurfSense fetches from live platforms and returns structured JSON that agents can consume. The README says the project has redirected its energy toward giving agents live data, while the knowledge base, cited chat, reports, podcasts, and presentations continue to work.

Official sources

  1. Issues
  2. MODSetter/SurfSense on GitHub
  3. Project website
  4. README
  5. Releases
For maintainers

Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/modsetter-surfsense.svg)](https://hysenlabs.com/projects/modsetter-surfsense)
Community notes

Community notes