# Kaelio/ktx: a context layer that teaches Claude Code and Codex how to query your warehouse

> ktx ingests warehouse metadata, BI models and wiki pages, then serves the result to AI agents over MCP and a CLI. It is for teams whose agents keep reinventing metric logic, not for anyone without a SQL warehouse.

**Kaelio/ktx** — ktx is an executable context layer for data and analytics agents 🐙 Allow Claude Code, Codex, or other AI agents to query analytical databases accurately and with full context of your company

- Repository: https://github.com/Kaelio/ktx
- Website: https://docs.kaelio.com/ktx
- Stars: 1,609 · Forks: 105
- Language: TypeScript
- License: Apache-2.0
- Published: 2026-09-10 · Updated: 2026-09-10 · Language: en
- Canonical page: https://hysenlabs.com/projects/kaelio-ktx

## The problem ktx was built around: agents that re-explore your warehouse on every question

The README states the failure mode plainly: general-purpose agents "re-explore your warehouse on every question, invent their own metric logic, and return numbers that don't match approved definitions." That is a specific complaint, and it is the one ktx is aimed at. An agent with a database connection and no context will sample tables, guess at join keys and write its own revenue calculation, and the result will differ from the one finance reports.

The second half of the argument is aimed at traditional semantic layers. ktx's README says they "demand constant manual upkeep and don't absorb the rest of your company's knowledge." Whether you accept that framing or not, the design consequence is clear: ktx tries to generate and maintain the semantic layer rather than asking a data team to author it by hand.

The audience follows from that. The README lists teams that want Claude Code, Codex, Cursor or OpenCode querying a warehouse with approved definitions, and teams whose business knowledge is scattered across dbt, Looker, Metabase, Notion and internal wikis. The README is equally direct about who should not bother: anyone without a SQL warehouse, since ktx sits on top of one, and anyone who needs a single ad-hoc query.

## How ktx builds context: connectors, reconciliation, then Markdown and YAML

The repository ships two diagrams that describe the pipeline better than the prose does. The ingestion diagram names four stages: source connectors, a context builder, reconciliation, and validation. The output of that pipeline is two artifacts, wiki Markdown and semantic-layer YAML. That is the shape of the system: ingestion is a batch process that produces files, not a live query proxy.

On the input side, the connectors cover databases (PostgreSQL, Snowflake, BigQuery, ClickHouse, MySQL, SQL Server, SQLite, DuckDB, Amazon Athena, MongoDB) and modelling or BI sources (dbt, MetricFlow, LookML, Looker, Metabase, Sigma, Notion, Google Drive). The README describes what the builder does with them: it samples tables, captures metadata and usage patterns, detects joinable columns, and annotates sources. It also ingests wiki content, organizes it, removes duplicates and flags contradictions for human review.

The serving diagram shows the other half. An agent queries ktx through MCP; ktx searches the wiki and the semantic layer, returns approved metrics, and compiles them into read-only SQL that runs against the warehouse. Two claims in the README are worth separating. The join graph is said to resolve chasm and fan traps automatically, which is the part a hand-maintained semantic layer normally handles. And the whole path is described as read-only by design, which is the constraint that makes it safe to point at a production warehouse.

Search at execution time is hybrid: the README says the CLI and MCP tools expose combined full-text and semantic search across wiki and semantic-layer entities. Embeddings are part of that, since `ktx status` reports an embeddings model alongside the LLM.

## Installing ktx and running a first real query

ktx is published on npm as `@kaelio/ktx`. The README's Quick Start is three commands: install globally, run setup, check status.

```bash
npm install -g @kaelio/ktx
ktx setup
ktx status
```

`ktx setup` is not a single prompt. According to the README it creates or resumes a local ktx project, configures providers and connections, builds context, and installs agent integration. Expect it to ask about your LLM and embeddings providers and about at least one database connection.

`ktx status` is the check that matters before you open an agent client. The README shows an example of its output:

```text
ktx project: /home/user/analytics
Project ready: yes
LLM ready: yes (claude-sonnet-4-6)
Embeddings ready: yes (text-embedding-3-small)
Databases configured: yes (warehouse)
Context sources configured: yes (dbt_main)
ktx context built: yes
Agent integration ready: yes (codex:project)
```

If `ktx status` prints a line like `ktx mcp start --project-dir ...`, the README says to run it before opening your agent client. That is the MCP server the agent talks to. Once it is up, the first useful commands are searches rather than queries: `ktx sl "revenue"` searches semantic sources and `ktx wiki "refund policy"` searches local wiki pages. If you would rather have the agent install itself, the README gives an alternative path: ask Claude Code, Codex, Cursor or OpenCode from the project directory to run `npx skills add Kaelio/ktx --skill ktx`.

One note on credentials: the README says ktx runs with your own LLM API keys or a local agent sign-in, either a Claude Pro/Max subscription through Claude Code or your local Codex authentication, and that there is no extra usage billing from ktx.

## Where ktx stops being the right tool

The README names two cases itself, and they are the honest ones. Without a SQL warehouse there is nothing for ktx to sit on. And for a single ad-hoc query, psql or a notebook is the correct answer; standing up connectors, embeddings and an MCP server to answer one question is wasted work.

There is a third boundary the README implies rather than states. Ingestion is a build step, and its inputs are other systems: dbt projects, LookML, BI tools, wiki pages. A team with none of those has little for the context builder to reconcile, and the join-graph and contradiction-flagging behaviour has nothing to work on. The value scales with how much scattered knowledge you already have, which means it is lowest exactly where setup is easiest.

The reconciliation step is also the part most likely to need a human. The README says contradictions between sources are flagged for review, which is a sensible design, but it means someone has to sit down and resolve them. ktx reduces manual upkeep of a semantic layer; it does not eliminate review. Treat the flagged contradictions as a queue you own.

Finally, the repository is a pnpm and uv workspace with a TypeScript CLI and Python packages for the semantic layer and a daemon. The root `package.json` sets `"node": ">=22.0.0"` and `"pnpm": ">=10.20.0"` as engine requirements, and the Python workspace requires 3.13 or newer. If your environment is pinned below those, that is a real constraint before you evaluate anything else.

## ktx against a hand-maintained semantic layer

The obvious alternative is the traditional semantic layer, and ktx's own comparison table makes the difference explicit. Both columns get a check for "approved, reusable metric definitions." The divergence is in the other rows. Detecting joinable columns and resolving fan and chasm traps is marked "Manual" for a traditional semantic layer and a check for ktx. Absorbing wiki, Notion or team knowledge is a dash for the traditional layer and a check for ktx.

That is the real trade. A hand-authored semantic layer is deliberate: a person decides what a metric means, and the definition is reviewable in a pull request. ktx infers much of that from sampled tables, usage patterns and ingested documents, then reconciles across sources. Inference is what makes it cheap to start and what makes its output worth auditing. If your organisation's metric definitions are politically sensitive, or already carefully maintained, generating them from samples is a downgrade in control even when it is an upgrade in coverage.

The other alternative is doing nothing and letting a general-purpose agent work directly against the warehouse. The comparison table gives that column a dash for building warehouse context automatically, a dash for absorbing company knowledge, and "Partial" for shipping a CLI and MCP for agent execution. The cost of doing nothing is not zero; it is the re-exploration and invented metric logic the README opens with.

## Licence, releases and what upgrading costs

ktx is Apache-2.0, and the licence file sits at the repository root. That is a permissive licence with an explicit patent grant, which matters if you are embedding the tool in a commercial data stack. The `pyproject.toml` for the Python workspace also declares `license = "Apache-2.0"`, so the licensing is consistent across the TypeScript and Python sides. Nothing here is legal advice; read the licence text before you rely on it.

The release cadence visible in the repository is fast and the version numbers are pre-1.0. v0.14.0 and v0.15.0 both landed on 2026-06-30, and v0.16.0 followed on 2026-07-03. The last push to the repository was on 2026-09-03. Pre-1.0 versioning with that cadence means minor releases can change behaviour, and the README's upgrade instruction reflects that: re-run the global install with the `@latest` tag.

```bash
npm install -g @kaelio/ktx@latest
```

There is no documented rollback procedure in the README, and no mention of a pinned-version install for reproducibility. If you need to hold a version, that is something to work out from the npm package itself rather than from the documentation. On the Python side, `pyproject.toml` pins the uv required version to 0.11.11, so a developer environment has at least one hard tooling constraint written down.

## Conclusion

Adopt ktx if your agents already query a warehouse and keep producing numbers that disagree with approved definitions, and you have dbt, Looker, Metabase or wiki content worth ingesting. Skip it if you have no SQL warehouse, or if you only ever run one-off queries where psql or a notebook is enough. Before committing, verify two things yourself: that your warehouse dialect is on the supported list, and that `ktx status` reports context built and agent integration ready in your own project.

## FAQ

### What is Kaelio/ktx and who is it for?

ktx is a self-improving context layer that teaches agents how to query your warehouse using approved metric definitions, joinable columns and business knowledge. It is aimed at teams that want Claude Code, Codex, Cursor or OpenCode to query a warehouse accurately, and at teams whose knowledge is spread across dbt, Looker, Metabase, Notion and internal wikis.

### How do I install ktx?

Install the npm package globally with `npm install -g @kaelio/ktx`, then run `ktx setup` to create or resume a project, configure providers and connections, build context and install agent integration. Run `ktx status` afterwards to confirm the project is ready.

### Which databases and tools does ktx work with?

The README lists PostgreSQL, Snowflake, BigQuery, ClickHouse, MySQL, SQL Server, SQLite, DuckDB, Amazon Athena and MongoDB as databases, and dbt, MetricFlow, LookML, Looker, Metabase, Sigma, Notion and Google Drive as integrations.

### Does ktx need its own LLM subscription?

No. The README states that ktx runs with your own LLM API keys or a local agent sign-in, either a Claude Pro/Max subscription through Claude Code or your local Codex authentication, and that there is no extra usage billing from ktx.

### Can ktx write to my warehouse?

The README describes the serving path as read-only by design, and the serving diagram shows approved metrics being compiled into read-only SQL that runs against the warehouse. Nothing in the README describes write access.

## Sources

- [Kaelio/ktx on GitHub](https://github.com/Kaelio/ktx)
- [License: Apache-2.0](https://github.com/Kaelio/ktx/blob/main/LICENSE)
- [Project website](https://docs.kaelio.com/ktx)
- [README](https://github.com/Kaelio/ktx/blob/main/README.md)
- [Releases](https://github.com/Kaelio/ktx/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/kaelio-ktx
