Model or dataset
pzqpzq/Principia avatar
pzqpzq/Principia

Principia's five demo results are measured against three different baselines, and one of them is the mean

Principia extracts reusable principles, composes those principles into traceable research ideas, and helps researchers inspect why an idea may be worth testing.

1,019 stars52 forksRich Text FormatApache-2.0

At a glance

What is it?
Principia is a local workspace that takes a research goal and a dataset, fits candidate expressions, challenges them with held-out checks, and stores the surviving ones as Rules with equations and a stated scope. The demo projects publish their error figures. Reading those figures closely shows a lot about what the tool claims: the loudest improvement is measured against a predictor that outputs the training mean, and one of the five recovered relationships is the Pythagorean identity.
Who is it for?
Read this as an evidence-recording tool rather than a discovery engine. The design decision worth copying is that a candidate must survive a challenge stage before it becomes a Rule, that failed candidates stay in the record, and that every Rule carries its own scope.
Can I use it commercially?
Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 2 days ago.
What is it written in?
Mainly Rich Text Format, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

Three demos, three baselines, and the mean is one of them

The demo table publishes a held-out error for each project, and the baselines are not the same object. The field calibration project reports held-out errors of 0.0825 and 0.0923 volts per metre for its optical and ion readouts, described as more than 92 percent below their baselines, and the word before baselines is development mean. Predicting the mean of your development set is the cheapest possible predictor, so a 92 percent reduction says the quantity moves enough to be worth a model and nothing more. The earthquake project compares against a frozen linear baseline instead, and the wafer project against a constant baseline, which is the same idea as the mean but with the word changed. So three of the five headline numbers are reductions against a predictor with no structure at all. The page does not hide this, since each row carries its own scope note, but the percentages read very differently once the comparison is named.

One of the five recovered relationships is an identity that holds by definition

The particle physics project is the clearest case. The rule it reports is a parameter free vector identity for transverse momentum, the square root of the sum of the two transverse components squared, with a direction check alongside it. The evaluation covers 78,227 events across two held-out period files, and the reported error is about 1.17 times ten to the minus seven at the 99th percentile. An error at that magnitude is floating point noise, which is what you get when a formula is true for every input rather than approximately true for a dataset. The page says as much in its own scope note, describing the result as checking representation consistency and recovering known geometry. Included in a demo set, this row is useful: it shows the evaluation harness can confirm a correct identity across a large file. Read as a discovery, it is a tautology, and the page is careful enough not to claim otherwise.

Held-out data is thin in two of the five projects

The scope notes are where this project is unusually disciplined, and they undercut two results. The calibration project has only two held-out power settings per readout, and both readouts share the same radio frequency chain, so the test covers two points along a curve the development data already traced. The earthquake project splits thirty one days of catalog into eighteen for development, six for validation, and seven for testing, restricted to a retained subset of the catalog, and the page states plainly that it does not predict event times or locations. The wafer result is a fit for the measured specimen with transfer to another wafer untested. The walking project is described as a consistency relationship between contact moment, normal load, and pressure centre displacement under a recorded coordinate convention, and it is explicitly not evidence of a new biological mechanism. Five demos, and three of them are demonstrations of arithmetic consistency rather than findings.

A fit does not become a Rule until the challenge stage lets it

The workflow table has four stages and the third one is the load bearing. The first stage reads the supported files, profiles the measurements, and connects variables to the research goal and the scientific context, with the inspectable output being an inventory of formats, units, coverage, and interpretation. The second proposes candidate expressions, fits them on development data, and uses validation evidence to compare them. The third applies recorded held-out checks and evidence gates before anything executable is promoted, and what you can inspect there is the split definition, the errors, the controls, the failure reasons, and the limits of applicability. Two sentences carry the philosophy. A strong fit alone does not establish a mechanism, and a dataset may produce observations without a Rule that passes the available checks. Unsuccessful candidates stay in the study record rather than being deleted.

Four object types, each carrying its own evidential scope

The design separates four things that most discovery tools blur. Literature Principles supply scientific context and carry a stated scope you can read without leaving the project. Observations describe computed evidence. Rules are the executable results, and each one carries an equation, calibration results, validation decisions, and its own scope. A study map connects the three so you can move from a promising relationship down to the records that support it. The page is explicit that known identities, empirical relationships, and tentative interpretations keep different evidential scopes rather than being promoted into one category, and that a Rule is rendered as a typeset equation with a plain language interpretation and links to the underlying numerical records, with rounded display values backed by full precision ones. That separation is the part of this workspace worth taking regardless of whether you keep the rest.

The demos open without the data or a key, so the evidence is readable and not reproducible

The five demo projects are public datasets analysed with one named model, and their saved study maps, results, equations, and packaged evidence open locally without the original datasets and without an API key. That is a deliberate choice and it cuts both ways. In your favour, you can read every claim the project makes without spending anything or trusting a hosted service. Against you, the thing you cannot do is re-run the analysis that produced the numbers, because the input files are not in the package and the model call is not reproducible without a key. The page is straightforward about the arrangement, telling you to connect your own data and models when you want to start a discovery, and pointing at a separate demo guide for portability, reruns, and deletion. The workspace also defaults to discovering again as a new project, keeps data and workspace state separate so the source dataset is not copied per project, and exposes task activity, stop controls, and project deletion.

The version number is a directory name, twice over

The repository root is a wrapper rather than the application. The product sits inside a directory named for its release, and inside that sits a second directory named for the core release, which contains its own licence file and its own documentation tree, with the demo guide reached through a path that repeats the version a third time. So a release is a new pair of directories rather than an edit to existing ones, and every path a script hardcodes changes on upgrade. Two of the root directories, one for a global cloud deployment and one for a scenario set, are not explained anywhere in the visible page. The benchmarks directory is linked from the header, so at least that one has an entrance. The licence exists at both levels, which is more than most projects do, though it means there are two files to keep in step.

Three same-day tags named after commit hashes, and a language detector reading RTF

The release list is unusual and worth a look before you decide how to track versions. All three tags carry the same date prefix, a global marker followed by a date in August 2026, and end in twelve hexadecimal characters that look like commit hashes, and all three were published on the same day. They are not semantic versions, so there is no ordering you can reason about, and the list is not even in time order, with the middle timestamp of the three attached to the last entry. Alongside that, the project records its primary language as Rich Text Format while its own setup section asks for Python 3.11 or 3.12 in a virtual environment. The most recent push was on 2026-10-02. None of this is a reason to avoid the tool; it is a reason to read the workspace files rather than the tag names when you want to know what you are running.

Editorial conclusion

Read this as an evidence-recording tool rather than a discovery engine. The design decision worth copying is that a candidate must survive a challenge stage before it becomes a Rule, that failed candidates stay in the record, and that every Rule carries its own scope. Before drawing conclusions from the demos, check what each baseline is: two of the five compare against a constant or mean predictor, and one recovered result is an identity that holds by definition. And since the packaged evidence opens without the original data, treat every number as a claim to re-derive rather than a result you can rerun from what ships.

Frequently asked questions

What does Principia do?

It takes a research goal and a local dataset, connects scientific context with executable analysis to look for compact interpretable relationships, and keeps the evidence behind each result available. A surviving relationship becomes a Rule carrying an equation, calibration results, validation decisions, and a stated scope.

How does Principia decide a relationship is worth keeping?

A challenge stage applies recorded held-out checks and evidence gates before an executable relationship is promoted. The page states that a strong fit alone does not establish a mechanism, and that a dataset may produce observations without a Rule that passes the available checks. Failed candidates stay in the study record.

What models does Principia use?

The five demo projects were analysed with DeepSeek-V4-Pro, and the workspace supports configurable reasoning and vision models. The page says to connect your own data and models when you want to begin a new discovery.

Do the Principia demo projects need an API key or the original data?

No. The saved study maps, results, equations, and packaged evidence open locally without the original datasets and without an API key, so the recorded claims can be inspected but the analysis that produced them cannot be re-run from what ships.

What baselines are the Principia demo results compared against?

Three different ones. The field calibration project compares against development mean baselines, the earthquake project against a frozen linear baseline, and the wafer project against a constant baseline.

Official sources

  1. Issues
  2. License: Apache-2.0
  3. pzqpzq/Principia on GitHub
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/pzqpzq-principia.svg)](https://hysenlabs.com/projects/pzqpzq-principia)