Open-source project
osintbrazuca/osint-brazuca avatar
osintbrazuca/osint-brazuca

OSINT Brazuca: a catalogue of Brazilian public data sources, indexed two ways

Repositório criado com intuito de reunir informações, fontes(websites/portais) e tricks de OSINT dentro do contexto Brasil.

2,779 stars370 forksPythonMIT

At a glance

What is it?
Two thousand seven hundred stars on a repository that is mostly a Markdown table of Brazilian government portals, plus a generated JSON dataset that turns that table into something searchable by what you already have and what you want back.
Who is it for?
The specific gap this repository fills is narrow and real: almost every general OSINT resource is written for United States or European data, so a Brazilian investigator knows the technique but not the URL. OSINT Brazuca supplies the URLs, and the generated JSON in `data/` supplies something a link dump usually lacks, which is a query path from an input you hold to the sources that accept it.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 7 days ago.
What is it written in?
Mainly Python, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A link catalogue that grew a machine-readable layer

The repository describes itself as a catalogue of Brazilian open sources, gathered to hold information, websites and portals, and techniques within a Brazilian context. The framing sentence in the README is that the point is to know these sources in order to protect, and to investigate in order to defend.

The interesting engineering decision is not the catalogue itself, it is the layer on top. Alongside the Markdown, the project publishes a structured JSON version under `data/`, generated automatically from the README. That generation step is what turns a link dump into something a script can use, and the README is explicit that it is derived rather than maintained by hand:

text
| [`data/sources.json`](data/sources.json) | 435 fontes em 40 categorias, com os links agrupados. Arquivo canônico. |
| [`data/index.json`](data/index.json) | 1.040 links achatados, um registro por URL. Formato pronto para busca. |
| [`data/taxonomy.json`](data/taxonomy.json) | Vocabulário controlado: 30 tipos de entrada, 38 de retorno e 14 de fonte. |
| [`data/overrides.json`](data/overrides.json) | Correções manuais de classificação. |

Four files with four distinct jobs. `sources.json` holds 435 sources across 40 categories with links grouped, and is named the canonical file. `index.json` is the same material flattened to 1,040 records, one per URL, described as ready for search. `taxonomy.json` is a controlled vocabulary of 30 input types, 38 return types and 14 source types. `overrides.json` holds manual corrections to classification, which is the escape valve that makes the rest trustworthy enough to automate.

The distinction between the first two is the one worth internalising. If you want the structure, read `sources.json`. If you have a lookup and want an answer, read `index.json`. The taxonomy counts are what make the catalogue useful rather than merely long, because they are what let a tool ask which sources accept a vehicle plate and which of those return an address.

Querying by input type rather than by category

Most link collections are organised by subject, so you scroll a table of contents until you find the right heading. This one is organised by data type instead, and the README states the axis clearly: you can locate sources by input type, which means what you already have in hand, such as a CPF, a CNPJ, a vehicle plate, a domain or a phone number, and by return type, meaning what data the source gives back.

That is a better shape for investigation because it matches how an actual task starts. You do not begin with a subject area, you begin with one identifier someone handed you. Both axes are small closed vocabularies in `taxonomy.json`, which is what makes a query like find-me-a-source-for-a-plate answerable without parsing prose.

The other point about the dataset is a warning about the generator. The README marks `sources.json` and `index.json` as generated, and `overrides.json` as the place where manual classification corrections live. If a classification looks wrong to you, that file is the sanctioned path, not a patch to the generated output, because the next regeneration would discard the patch.

One thing to settle before planning around this: GitHub reports the primary language of the repository as Python, yet the top-level tree is Markdown, JSON, an `assets/` directory and a `tools/` directory with no visible Python at the root. The scripts that produce and validate the dataset are therefore in `tools/`, which is consistent with the README describing the JSON as generated automatically from the Markdown. If you want to regenerate rather than consume the data, that directory is where to start.

The legal framing is Brazilian law, not a disclaimer

The README spends more space on legal and ethical framing than on technique, which is the correct allocation for this kind of project and also the part that carries real information. It is not boilerplate about being a security tool. It names specific statutes and attaches obligations to each one.

The three cited are Lei nº 13.709/2018, the LGPD data protection law, Lei nº 12.965/2014, the Marco Civil that governs internet framework in Brazil, and Lei nº 12.527/2011, the access to information law. The README's instruction is that every query must respect the LGPD, and that processing personal data needs a legal basis and a legitimate purpose. That is a stronger claim than a warning label, because it says data being public does not by itself make processing lawful.

The permitted practices list is concrete: use only official public sources, respect people's privacy and dignity, document sources and methodology, have a legitimate purpose such as journalism, research, security or compliance, and do not publish sensitive data.

text
- Utilizar apenas fontes públicas oficiais
- Respeitar a privacidade e dignidade das pessoas
- Documentar fontes e metodologia utilizada
- Ter propósito legítimo (jornalismo, pesquisa, segurança, compliance)
- Não compartilhar dados sensíveis publicamente

The prohibited list is equally specific and worth reading as a scope definition rather than a warning: social engineering or system intrusion, stalking or harassment, use for discrimination or prejudice, unauthorized sale of data, and breach of professional secrecy.

For a non-Brazilian reader the mechanism is still useful even if the statutes are not binding, because it models how to scope a public-data project so that a compliance officer can read the README and find the reasoning already written down. The repository is MIT licensed, 2,736 stars, 363 forks, and only 2 open issues, with the last push on 2026-09-18 and no archive flag.

Operational limits the README states instead of leaving you to find out

There is a section on limitations that is more useful than most projects' troubleshooting notes, because the constraints it lists are properties of Brazilian government portals rather than bugs anyone intends to fix.

On data freshness: public data may be out of date, you should always check the last update date on the sources themselves, you should cross-check information from multiple sources for validation, and government systems may be down for maintenance. Each of those four is a distinct failure mode, and the first one is the one that will actually mislead you, because a stale record looks exactly as authoritative as a current one.

On access: some portals require prior registration through gov.br, services may have usage limits, and CAPTCHA may restrict automated queries. That third point is the practical constraint on automation, and it means any scraper built on this catalogue has to handle the case where it simply cannot ask.

On automation specifically, the README asks for respect for request limits, use of caching to reduce requests, and warns that bulk queries may get blocked:

text
- APIs públicas possuem **limites de requisições**
- Respeite os **limites técnicos** estabelecidos
- Use **cache** quando possível para reduzir requisições
- Consultas em massa podem ser **bloqueadas**

The final block of the README's warnings is the one that generalises furthest. It says information is public but protected by the LGPD, improper use can result in legal and criminal sanctions, sensitive data should not be shared publicly, the legitimate purpose of every query should be documented, a record should be kept of all research performed, and it closes with the informal instruction not to be a jerk. Read as a methodology note, the list is a reasonable minimum for any public-source investigation regardless of jurisdiction: state the purpose, log what you did, and do not republish what you found.

Four companion documents and a maintainer-facing contribution guide

The navigation badges at the top of the README point at five companion documents, and the split between them tells you what kind of reader each is aimed at.

text
README Principal
Exemplos Práticos
Fluxogramas
Guia Rápido
Contribuir

n `GUIA_RAPIDO.md` is the quick guide, `EXEMPLOS_PRATICOS.md` holds worked examples, and `FLUXOGRAMA.md` holds flowcharts, which is the right way to express branching investigation logic in Markdown. `CONTRIBUICAO.md` is the contribution guide, and the near-empty issue tracker with 2 open issues against 2,736 stars says a lot about how contributions actually flow: people propose sources rather than report bugs.

The README names two authors, Cleiton P. (MrCl0wnLab) and Diego (c4nh0t0), with GitHub and social profiles for each. Two maintainers for a project of this size is worth noting, since a catalogue whose value depends on links staying alive is exposed to a single person's availability.

The repository topics are brasil, hacking, osint, threat-hunting, threat-intelligence and threatintel, which is an honest signal about how the catalogue gets used even though the README's stated framing is defensive. There is no release history in the repository at all, so there is no version to pin and no changelog to read, which is the expected shape for a documentation project.

What this catalogue is not

It is worth being blunt about the boundaries, because the name invites the wrong expectations. This is not a tool. There is no scanner, no query engine, no enrichment pipeline, and nothing that takes an identifier and returns a result on its own. What it contains is a map of where Brazilian public data lives, plus a machine-readable index of that map.

That distinction matters for how you budget effort. Building on the catalogue means you still own every hard part: handling CAPTCHA on portals that require one, managing session and rate limits, normalising inconsistent formats across 40 categories, and verifying that a record is current. The catalogue saves you the part where you guess which portal might hold a given dataset, which is genuine work, but it is bounded work.

It is also a Brazilian catalogue specifically. Nothing here helps with an equivalent question in another jurisdiction, and there is no equivalent in this repository of internationally scoped sources, which means you would use a general OSINT resource for those and this one for Brazilian identifiers. The taxonomy of 30 input types and 38 return types is the concrete boundary of coverage, and checking it against the identifier type you actually hold is the fastest way to know whether this repository helps at all.

For the case it was built for, a Brazilian threat analyst, fraud investigator, journalist or security researcher, the value is concentrated exactly where generic OSINT material is thinnest: the specific URLs, and the legal context that governs using them.

Editorial conclusion

The specific gap this repository fills is narrow and real: almost every general OSINT resource is written for United States or European data, so a Brazilian investigator knows the technique but not the URL. OSINT Brazuca supplies the URLs, and the generated JSON in `data/` supplies something a link dump usually lacks, which is a query path from an input you hold to the sources that accept it. What it does not supply is the substance behind the links, so treat every entry as a starting point whose current validity you confirm yourself, exactly as the README advises about stale public data. Start with `data/index.json` if you are scripting, `GUIA_RAPIDO.md` if you are learning, and `data/overrides.json` if you find a classification that is wrong.

Frequently asked questions

Is OSINT Brazuca a tool or a list of links?

It is a catalogue, not a scanner. The repository holds Brazilian public data sources and publishes them as both Markdown and a generated JSON dataset under data/, but it does not take an identifier and return a result. You still handle portals, rate limits, CAPTCHA and data freshness yourself.

How do I find a source for a specific identifier such as a CNPJ or a vehicle plate?

The generated dataset is organised by input type, which is what you already hold, and by return type, which is what the source gives back. data/index.json holds the flattened records, one per URL, described as ready for search, while data/sources.json groups the same 435 sources into 40 categories and is the canonical file.

What laws does the project say apply to using Brazilian public data?

The README names Lei nº 13.709/2018, the LGPD, alongside Lei nº 12.965/2014, the Marco Civil, and Lei nº 12.527/2011, the access to information law. Its position is that data being public does not make processing it lawful: personal data handling needs a legal basis and a legitimate purpose.

Why does an entry in the catalogue sometimes turn out to be wrong or out of date?

Because the JSON is generated from the Markdown catalogue and many Brazilian government portals change or go down for maintenance. The README warns that public data may be out of date and that you should cross-check multiple sources, and data/overrides.json is the sanctioned place to record a manual classification correction rather than editing generated output.

Official sources

  1. Issues
  2. License: MIT
  3. osintbrazuca/osint-brazuca on GitHub
  4. README
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/osintbrazuca-osint-brazuca.svg)](https://hysenlabs.com/projects/osintbrazuca-osint-brazuca)