# google-research: a dump of released code, told not to be cloned

> The google-research repository is a flat archive of research project folders with two licenses, no releases, and instructions that ask you to take one directory and leave. Here is what that arrangement costs the reader who wants the code.

**google-research/google-research** — Google Research

- Repository: https://github.com/google-research/google-research
- Website: https://research.google
- Stars: 38,849 · Forks: 8,475
- Language: Jupyter Notebook
- License: Apache-2.0
- Published: 2026-08-17 · Updated: 2026-08-18 · Language: en
- Canonical page: https://hysenlabs.com/projects/google-research-google-research

## The first instruction is not to clone the repository

The opening request of this repository is restraint. Because the repo is large, the advice is to download only the subdirectory of interest, and the route given is the web editor: change the address from github.com to github.dev, then right-click the folder of interest in the left navigation panel and select download.

Read that as a policy rather than a convenience. The project is a place where released research code is put on display, not something you install, and the person who wrote those lines expected most visitors to take a single folder and leave. The consequence for a reader is that you get a directory of scripts and notebooks with nothing around it, no version marker beyond whatever state the folder carried that day, and no way to tell from the download where it came from. If you plan to keep the code, write down the date you took it.

## The one command given is an SSH shallow clone

Pull requests are the exception, and they come with a single line:

```bash
git clone git@github.com:google-research/google-research.git --depth=1
```

Two details of that line matter. It uses the git@github.com SSH form, so it needs a GitHub SSH key on the machine before it will run at all, and no HTTPS alternative is offered for a reader who has only a token. And the --depth=1 flag is the point of the recommendation, because dropping the history is what keeps a repository this size workable.

The consequence lands on anyone who contributes. A shallow clone has no history, so git log, blame, and bisect stop answering the moment a review question turns to where a line came from. The README makes that bandwidth trade on your behalf, and it is the wrong trade for a long argument about provenance.

## Apache 2.0 for source, CC BY 4.0 for data, and no map between them

Licensing here splits by file type rather than by directory. All source files are released under the Apache 2.0 license, with the text in the LICENSE file, and all datasets are released under CC BY 4.0 International, whose legal code lives at creativecommons.org/licenses/by/4.0/legalcode.

Since every project sits in its own top level folder, and a project folder can hold both scripts and data, the folder name tells you nothing about which terms apply to what is inside it. No per project license listing exists, and nothing maps paths to licenses. The consequence is a compliance task done by hand: inspect the file you are actually copying, keep the Apache notices with the code, and carry the attribution that CC BY 4.0 asks for, which is an obligation a permissive code license does not impose.

## The instructions are stamped 2023, the last push is 2026-09-23

The file ends with two lines that matter more than they look. The first is a disclaimer: this is not an official Google product. The second is a date, Updated in 2023.

That stamp and the repository history do not line up in a useful way. The last push to the default branch, master, is dated 2026-09-23, so code kept landing for years after the guidance was written. And the repository has published no releases, so there is no tag to install and no changelog to read. The consequence: the instructions are three years older than the code they point at, and the only coordinates you can record are a commit hash. Note the branch name as well, master rather than main, since tooling that assumes main will look at the wrong ref.

## The top level is a flat directory list with no catalog

The top level holds LICENSE, README.md, CONTRIBUTING.md, an __init__.py, a .circleci directory, a .gitignore, and then a long alphabetical run of project folders. Some are short acronyms: aqt, alx, albert, CoDi, TimesX, STraTA, OpenMSD. Others are whole paper titles with the spaces turned into underscores, such as CIQA, LLP_Bench, RevThink, and the long names about learning linear thresholds from label proportions.

So the tree mixes two naming conventions with no index, no table of contents, and no per directory summary anywhere in the README. A reader hunting for a specific paper has to guess which folder holds it, and a reader inspecting a folder has to guess which paper it came from. Nothing connects a directory name to a citation, a dataset, or an author list, so a folder called TimesX or aav explains nothing on its own.

## A notebook repository with no dependency manifest

The primary language of this repository is Jupyter Notebook, which tells you what a project folder is: something you execute, not a package you import. There is no install command in the README, and no dependency manifest is named there either.

The one Python shaped file at the root is __init__.py, which makes the whole tree importable as a single package with every project as a top level name inside it. Two things follow from that layout. Every project name is unique by construction, so a clash between two of them would have to be settled by renaming a directory. And the import root is the repository itself, so the projects share one namespace rather than living under a vendor prefix. For a reader, the practical effect is that a notebook's dependencies are whatever the notebooks assume, and you are handed no list to check them against.

## Pull requests are invited and their rules are not written down

The README says that submitting a pull request requires a clone and recommends the shallow one, but it does not point you at CONTRIBUTING.md even though that file sits at the top level. The only continuous integration signal in the tree is a .circleci directory, and no step, lint rule, or test command is written down anywhere.

The consequence is a contributor workflow that resolves after the push instead of before it. You learn what CI enforces, and whether a shallow clone is acceptable, from the check run rather than from documentation. For a reader who only wants one folder, the same gap appears differently: nothing in the repository states which Python version, which accelerators, or which dataset layout a given project expects, so compatibility is something you establish per directory, on your own machine, one folder at a time.

## Conclusion

This repository fits a reader who wants one specific paper's code and can wire up its dependencies by hand, and it fits a contributor who can work inside a shallow clone. It is a poor fit if you need a versioned dependency, a license you can read from a directory name, or onboarding rules written down before you push. Before you build on a folder, record the commit you downloaded, check the license of the specific files you copy, and confirm the environment the notebooks expect on your own machine.

## FAQ

### what is google research

The repository is a public dump of code released by Google Research, with source files under Apache 2.0 and datasets under CC BY 4.0 International. It carries a disclaimer that it is not an official Google product, a note that the file was updated in 2023, and a homepage at research.google.

### is google research ai

The top level is a flat alphabetical list of research project folders, from short names such as aqt, alx and TimesX to full paper titles with spaces replaced by underscores, alongside a root __init__.py. Nothing groups those folders or states which of them are current, and no GitHub releases exist for the repository.

### what is google research colab

Nothing in the repository mentions Colab. The closest documented route to running a single project is to open it in the GitHub web editor by changing github.com to github.dev and downloading that folder, or to clone the whole repository shallowly with git clone git@github.com:google-research/google-research.git --depth=1.

### How do I get into Google Research?

The repository documents no hiring, internship, or onboarding route. The only way in it describes is a pull request, which requires cloning the repository, and it recommends a shallow clone without history to keep the download manageable.

### google research vs deepmind

The repository contains no comparison and never mentions DeepMind. What it does state is that it holds code released by Google Research, that it is not an official Google product, and that its default branch is master rather than main.

## Sources

- [Official documentation](https://research.google)
- [Official README](https://github.com/google-research/google-research#readme)
- [Project repository](https://github.com/google-research/google-research)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/google-research-google-research
