# ScienceClaw: Self-Evolving AI Research Colleague for Scientists

> ScienceClaw is a TypeScript research system built on the OpenClaw engine that starts with 285 skills and writes new ones during sessions, stores research patterns across weeks using LanceDB, and enforces zero citation fabrication through a 629-line protocol. The design pays off for researchers running extended multi-database literature reviews; it is not a substitute for a general-purpose AI assistant.

**beita6969/ScienceClaw** — 🔬🦞 A self-evolving AI research colleague for scientists. 285 skills, zero hallucination, persistent memory.

- Repository: https://github.com/beita6969/ScienceClaw
- Website: http://scienceclaw.science
- Stars: 907 · Forks: 105
- Language: TypeScript
- License: MIT
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/beita6969-scienceclaw

## What ScienceClaw Solves: Multi-Database Research With Verifiable Citations

General AI assistants end their memory when a conversation closes, apply no minimum depth requirements, and may generate plausible-looking citations for papers that do not exist. Researchers who need to pull from PubMed, Semantic Scholar, NCBI Entrez, and domain-specific databases in one extended session have no built-in assurance that a cited paper is real. ScienceClaw addresses this by building on the OpenClaw engine (described in the package.json as a multi-channel AI gateway) and layering a research-specific design on top: domain skills for natural and social sciences, persistent memory across sessions, and a hard protocol requiring that every citation can be traced to a tool result in the current conversation.

The project targets researchers who run literature reviews, not engineers looking for a coding assistant or a general chatbot. Its skill coverage spans biomedicine (PubMed, UniProt, KEGG, PDB, ClinicalTrials, gnomAD), chemistry (PubChem, ChEMBL, RDKit), genomics (NCBI Entrez, Ensembl, ClinVar, GEO), materials science (Materials Project, pymatgen), physics (astropy, quantum-computing, simulation), and environmental science (Copernicus climate data, GIS tools). Social sciences have their own skill sets. This breadth distinguishes it from tools aimed only at biomedical literature.

## How the Self-Evolving Skill System Grows With Your Research

ScienceClaw starts with 285 skills. Each skill is stored as a SKILL.md file. When a research task completes, the system writes new SKILL.md files at runtime without any redeployment. A skill might record that PubMed combined with Semantic Scholar yields better recall for immunology queries than either source alone, or that a specific output format with PMID and DOI is always required. Over weeks of use, the agent accumulates specialized search templates, database priority chains, and output preferences tuned to a subfield.

The base OpenClaw engine ships with roughly 54 general-purpose skills that do not change. Those skills serve messaging integrations and general tasks. They include no mechanism for learning a researcher's domain preferences. ScienceClaw's skill growth is persistent: skills written during one session are available in the next. A fresh installation starts from the 285-skill base; everything accumulated after that is specific to the user's research history.

The repository does not document whether skills learned for one user are isolated from others in a multi-user deployment. Teams considering shared deployments should treat that as an open question before going to production.

## Research Memory Across Sessions: LanceDB, Temporal Decay, and Pattern Retrieval

Standard AI conversation context resets when a session ends. ScienceClaw uses LanceDB for vector storage and adds two behaviors specific to long-running research: temporal decay weighting and cross-session research pattern retrieval.

Temporal decay weighting means that more recent findings carry more weight when the system reconstructs context for a new session. Cross-session pattern retrieval means you can open a new session and ask the agent to reuse the search strategy that worked for a previous project; the system retrieves the stored pattern and applies it.

Context window management differs from standard truncation. When the window fills during a long session, ScienceClaw applies smart compaction: it preserves statistical results, effect sizes, and key citations while discarding intermediate steps. The goal is to retain the research output, not the scaffolding, when the session must shed content.

The base OpenClaw engine includes a basic memory plugin. It stores some state but does not add temporal decay weighting, domain-aware compaction, or cross-session pattern retrieval.

## Configuration and First Run via setup.sh and the .env File

The repository contains a setup.sh script at the root. Running it after cloning handles prerequisites, environment setup, and dependency installation. The runtime reads configuration from a .env file at the project root or from ~/.openclaw/.env for daemon deployments.

The .env.example file maps the required variables. The gateway needs an authentication token, which the .env.example suggests generating with:

```bash
openssl rand -hex 32
```

Set the output as the value of OPENCLAW_GATEWAY_TOKEN. At least one model provider key is required. The .env.example lists the following as options:

```bash
# OPENAI_API_KEY=sk-...
# ANTHROPIC_API_KEY=sk-ant-...
# GEMINI_API_KEY=...
# OPENROUTER_API_KEY=sk-or-...
```

The README does not describe the full setup sequence in the portion that is available. The presence of setup.sh suggests it handles most of the steps automatically. The SCIENCE.md and SKILL.md files in the skills/ directory are configuration artifacts, not code to edit by hand.

## The SCIENCE.md Protocol: Why Citation Accuracy Takes Priority Over Speed

ScienceClaw's approach to hallucination is structural, not advisory. SCIENCE.md is a 629-line file that governs all agent behavior. The README describes it as the highest-priority rule in the entire system, one that applies before any other instruction.

The rule is specific: every citation must come from a tool result in the current conversation. If a database did not return a paper, the agent cannot cite it. If a result cannot be confirmed, the agent must say 'not verified' explicitly. If no evidence was found, it must say so rather than synthesize from general knowledge.

The practical difference from general-purpose language models is concrete. A model without tool-result citation requirements can produce a plausible DOI for a paper that does not exist. In scientific manuscript preparation, a fabricated citation discovered during peer review wastes reviewer time and damages the submitting author's credibility. ScienceClaw accepts slower response times as the price of source-verifiable output.

The protocol also prevents specific language: no 'I think,' no 'probably,' no hallucinated PMIDs. The README treats this as non-negotiable, which means a researcher cannot override it through a prompt instruction.

## Mandatory Depth Thresholds and One-Hour Sessions

ScienceClaw enforces minimum work before it concludes any task. Four task types carry minimum tool-call thresholds: Quick tasks require at least 5 tool calls; Survey tasks require 30; Review tasks require 60; Systematic reviews require 100 or more. Before concluding, the agent must have searched at least three different databases, retrieved full metadata rather than just titles, cross-referenced findings across sources, checked for contradictory evidence, verified key statistics against primary sources, and organized results into a structured output file. Any unmet condition causes the agent to keep working.

Session timeout reflects this design. The base OpenClaw engine has a default timeout of 600 seconds (10 minutes). ScienceClaw extends this to 3600 seconds (one hour), with a heartbeat that keeps sessions alive across interruptions. Querying PubMed, NCBI Entrez, Semantic Scholar, and ClinVar for a systematic review does not finish in ten minutes.

The mandatory depth enforcement has a trade-off: it cannot be switched off for quick checks. The README does not describe a mode that bypasses the minimum tool-call thresholds for simpler queries.

## Scope, Coverage Gaps, and When to Use OpenClaw Instead

ScienceClaw covers natural sciences and social sciences through separate skill sets. The database integrations span biomedicine, chemistry, genomics, materials science, physics, and environmental science on the natural science side, with social science disciplines covered separately. This scope is deliberately broad across academic fields.

The flip side is that ScienceClaw is not a general-purpose tool. It is not a coding assistant, a document editor, or a business workflow tool. The OpenClaw engine it extends was built as a multi-channel messaging gateway. ScienceClaw redirects that engine entirely toward academic research. A developer who needs coding assistance or a team managing business communications should use OpenClaw or a general-purpose AI assistant directly rather than running the research-specialized overlay.

Coverage also depends on connected databases. The no-hallucination rule means the agent can only cite what a tool returns in the current session. A paper that exists but is not indexed in any database the agent has access to cannot appear in the output. The quality of citations is bounded by the coverage of the connected data sources, not by what the underlying language model was trained on.

The last push to the main branch was on 2026-06-08. The project is MIT-licensed. The repository has no GitHub releases; deployment is from the main branch.

## Conclusion

ScienceClaw is a good fit for researchers who run systematic literature reviews across multiple databases and need verifiable citations from every tool call. It is not a fit for developers needing a coding assistant or teams looking for general business automation. Before deploying, verify that the databases your research requires are among the supported integrations, that at least one model provider API key is available, and that your environment can sustain one-hour sessions without a forced timeout.

## FAQ

### What distinguishes ScienceClaw from the OpenClaw engine it is built on?

ScienceClaw adds 285+ self-evolving research skills, LanceDB persistent memory with cross-session retrieval and temporal decay weighting, the SCIENCE.md citation protocol, and session timeouts of up to one hour. The base OpenClaw engine ships with roughly 54 general-purpose skills and a 10-minute default timeout.

### Which AI model providers does ScienceClaw support?

The .env.example lists OPENAI_API_KEY, ANTHROPIC_API_KEY, GEMINI_API_KEY, and OPENROUTER_API_KEY as options. Only one provider key needs to be configured; the project does not require all of them.

### Does ScienceClaw handle social science research or only natural sciences?

The README covers both natural sciences (biomedicine, chemistry, genomics, materials science, physics, environmental science) and social sciences through separate skill sets and database integrations.

## Sources

- [beita6969/ScienceClaw on GitHub](https://github.com/beita6969/ScienceClaw)
- [Issues](https://github.com/beita6969/ScienceClaw/issues)
- [License: MIT](https://github.com/beita6969/ScienceClaw/blob/main/LICENSE)
- [Project website](http://scienceclaw.science)
- [README](https://github.com/beita6969/ScienceClaw/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/beita6969-scienceclaw
