Model or dataset
alumnium-hq/alumnium avatar
alumnium-hq/alumnium

The Alumnium install puts your OpenAI key on a command line

End-to-End Testing with AI for Agents and Engineers

1,013 stars101 forksTypeScriptMIT

At a glance

What is it?
An AI-native testing library and MCP server that wraps the driver you already have, where the MCP registration passes an API key as a flag value, the Java quick start is a main method rather than a test, and the container exposes one port while starting the server on another.
Who is it for?
Alumnium is worth a trial on a flow that traditional selectors make brittle, because a check written as an intent sentence survives a markup change that would break a locator. Three things to settle first.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository received new commits within the last day.
What is it written in?
Mainly TypeScript, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 3, 2026, and from our analysis. They are not legal advice.

Editorial analysis

The MCP install instruction puts the key in a config file

Installing the MCP server is one script, and registering it with an agent is one line:

sh
curl -LsSf https://alumnium.ai/install.sh | sh
sh
claude mcp add alumnium --env OPENAI_API_KEY=... -- alumnium mcp

The shape of that second line is the problem. Passing a secret as the value of an environment flag to a registration command means the secret is an argument the tool records, and agent clients keep their server registrations, including environment values, in a configuration file on disk. The install path and the key management are therefore the same decision, and the documentation chooses convenience for it.

It is worth noticing what that implies about the tool generally. Every path in this project expects a key. The Python snippet sets an environment variable before constructing the driver, the TypeScript snippet assigns to the process environment first, and the Java snippet is the only one that does not show it at all, which reads as an omission rather than a difference. There is no local-model option in anything visible here, so the cost of running these tests includes a network call per check.

The Java quick start is a main method, not a test

The Java install declares test-scoped dependencies, and the contributor links at the bottom of the file point at a real test directory under the Java package. Then the quick start is this:

java
class AlumniumTest {
    public static void main(String...args) {
        ChromeDriver driver = new ChromeDriver();
        Alumni al = new Alumni(driver);
        al.act("type 'selenium' into the search field, then press 'Enter'");
        al.check("page title contains selenium");
        al.check("search results contain selenium.dev");
        al.quit();
    }
}

It is a plain entry point. There is no test annotation, no assertion, and no runner, in a snippet whose class is named like a test and whose dependency was added under test scope. The class name is doing the work the framework would normally do.

The other two languages go further than Java rather than matching it. Python calls a differently named method for the same action, uses do instead of act, then reads a value and asserts on it. TypeScript does the same with awaits. So the three supposedly equivalent examples cover three different amounts of the API, and the one whose ecosystem has the strictest conventions is the one that stops shortest.

The Java path pulls a per-platform binary with a version in the snippet

The Java dependency block is three lines and the second one is the interesting one:

groovy
dependencies {
  testImplementation 'ai.alumnium:alumnium:0.23.1'
  testRuntimeOnly    'ai.alumnium:alumnium-cli-darwin-arm64:0.23.1'
  // Add other platforms as needed
}

Two differences from the Python and TypeScript installs, which are one line each and resolve a wheel or a package. First, Java needs a separate command line artifact per platform, so a developer on Intel macOS or on Linux adds another line and a developer on Windows adds another. Second, the version is written into the snippet as a literal rather than read from a property, which means every upgrade is an edit to a build file by hand and the documentation's version can drift from the one in use.

The comment inviting you to add other platforms as needed is doing real work here, and it is the kind of instruction that ages badly. Nothing in the visible text pins a convention for keeping those lines in step with each other or with the library version, and the release numbers move often: three versions in the last month, the newest published within a second of the last push.

The TypeScript package is where the binary actually comes from, since the container build copies its output into place. Java is consuming a build artefact of a different language package, which is an unusual arrangement and explains why the version has to be written down twice.

The image exposes one port and starts the server without naming it

The container is small and careful in several places. It creates a dedicated unprivileged user with no log-init shell, chowns a cache directory to it, runs under a downloaded init binary, declares the cache as a volume, and switches to that user before the entrypoint. Those are all the right decisions for something that will execute test steps.

The port is where it slips:

dockerfile
EXPOSE 8013
VOLUME ["/app/.alumnium/cache"]

USER alumnium

ENTRYPOINT ["/tini", "--"]
CMD ["/app/alumnium", "server", "--host", "0.0.0.0"]

Nothing in the default command names a port. The image declares 8013, and the command passes a host of all interfaces and no port at all, so what the container actually listens on is whatever the binary defaults to, and the declared port is an assumption written in a different file. Anyone deploying this has to check which of the two numbers is real before they write a health check or a service definition.

The architecture handling next to it has the same shape of problem. The build stage picks a binary name by asking whether the target architecture is amd64 and treating every other answer as arm64, so any other target silently receives the wrong file.

The init binary is fetched during the build with no checksum

Two lines of the image are worth reading closely, and they concern that init process:

dockerfile
ADD https://github.com/krallin/tini/releases/download/v0.19.0/tini-${TARGETARCH} /tini
RUN chmod +x /tini

A build-time fetch of a release asset over HTTPS, pinned to a version number in the URL, with no checksum or signature verification step anywhere after it. The version pin means a changed asset at the same tag would be picked up silently, and a compromised release account would be picked up silently too. Nothing else in the image compensates, because this binary runs as the entrypoint, before the application and before the unprivileged user switch.

This is a small file doing a small job, and it is worth weighing only because of where it sits. Everything else in the image is either built from the repository or copied from a build stage inside it. This one piece of the runtime is defined by a URL in the Dockerfile.

The same pattern shows up in the first line of the install instructions, where a script is fetched and piped into a shell, and in the Java dependency version that is written by hand. Across all three, the choice is speed of setup against a verification step, and the project has consistently chosen setup.

A patch pinned to one exact dependency version, and four floors

The root manifest carries two mechanisms that most projects have one of. There is a patch map, keyed by an exact version:

json
  "overrides": {
    "ajv": ">=8.18.0",
    "lodash": ">=4.18.1",
    "picomatch": ">=2.3.2",
    "yaml": ">=2.8.3"
  },
  "patchedDependencies": {
    "@ai-sdk/[email protected]": "patches/@ai-sdk%[email protected]"
  }

The patch is the fragile half. Because the key names one exact version of one package, the patch applies to that release and nothing else. When the dependency moves to a new version, the patch quietly stops being applied and the fix it carried disappears, with no error, because an unused patch entry is a valid configuration. A patch that outlives its target is worse than no patch, since it looks like a fix in the repository.

The four floors are the safer half but not unconditionally. An override with a minimum version either pulls a transitive dependency up to satisfy the floor, or it fails to resolve, depending on whether any published release meets it. Both outcomes are correct behaviour; they just need checking against what actually exists on the registry before you assume a green install.

The root workspace has two scripts and a nightly compiler

The root manifest is marked private and declares two workspaces, the packages directory and a documentation site. Its script section has two entries, both of which invoke the compiler in project mode and one of which watches. There is no build script, no test script and no lint script at the root, so every one of those runs from inside a package or from a tool that is not wired here.

The dependency list explains the rest. There is exactly one runtime dependency at the root and it is unrelated to testing. The development side is where the project lives: the native preview of the TypeScript compiler, pinned to a dated build rather than a release, alongside the released compiler line at a major version above 5, plus a formatter, a linter, and a linter integration that bridges the native compiler's type information.

Relying on a dated nightly of a compiler that has not shipped a stable major is a deliberate bet on speed of type checking, and it means the toolchain can change under you without a version bump in this repository. Around it sit two task runners and an environment manager, with a lock file and an example local configuration, plus a workspace file and configuration for four or five different agent tools, including a root file that registers this project's own MCP server for contributors.

Editorial conclusion

Alumnium is worth a trial on a flow that traditional selectors make brittle, because a check written as an intent sentence survives a markup change that would break a locator. Three things to settle first. Decide how the key reaches the tool, since the documented MCP registration passes it as a flag value where it will be stored in your agent's configuration file, and every other path expects an OpenAI key as well. Treat the Java story separately from the Python and TypeScript ones, because Java pulls a per-platform binary with a version written into the snippet while the others are ordinary package installs, and the quick start you are shown for Java is a main method rather than a test. And if you deploy the container, check the port yourself, because the image declares one and the default command starts the server without a port argument while binding to every interface. If your requirement is reproducible, assertion-based tests with stable locators and no model in the loop, this is the wrong tool and you already know it.

Frequently asked questions

What is Alumnium?

An AI-native library and MCP server for end-to-end testing. It wraps a driver you already use, Appium, Maestro, Playwright or Selenium, and adds calls that act on a page and check a condition in plain language, plus a value read. It is published for Java, Python and TypeScript, and the MCP server installs with one script.

Which languages does Alumnium support?

Java, Python and TypeScript, each with its own install snippet and its own API naming. Python installs with pip and TypeScript with npm, while the Java build adds the library under test scope and a command line artifact for one platform, with a comment to add others.

Does Alumnium need an OpenAI API key?

Yes. The MCP registration command passes an environment variable holding an OpenAI key when adding the server to an agent, and the Python and TypeScript quick starts both set the same variable before building a driver. A configuration page is linked for the remaining settings.

How does the Alumnium container run?

From a Debian bookworm slim base, as a dedicated unprivileged user with a declared cache volume and a downloaded init binary as the entrypoint. The image exposes port 8013, while its default command starts the server with a host of all interfaces and no port argument, so the listening port depends on the binary's own default.

Official sources

  1. alumnium-hq/alumnium on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/alumnium-hq-alumnium.svg)](https://hysenlabs.com/projects/alumnium-hq-alumnium)