# tdxrs measures its own connection pool losing twelve times ground at 60 threads

> jiangtaovan/tdxrs is a Rust rewrite of the TDX market data parser with Python bindings, claiming 9 to 11 times on local file parsing. Its own tables say more about the design: pooled and async clients get twelve times slower at 60 threads while a per-request client stays flat, its rate limiter is fastest when the market is closed, and three parts of the documentation count the parser layer three different ways.

**jiangtaovan/tdxrs** — tdxrs 是通达信 (TDX) 行情数据解析库的 Rust 高性能实现，通过 PyO3/maturin 提供原生 Python 接口。它无缝兼容 [tdxpy] 的 API，并将核心解析引擎以 Rust 重构，从而实现数量级的本地解析性能提升，尤其在海量历史数据处理场景下优势显著。

- Repository: https://github.com/jiangtaovan/tdxrs
- Stars: 475 · Forks: 128
- Language: Rust
- License: MIT
- Published: 2026-09-20 · Updated: 2026-09-20 · Language: en
- Canonical page: https://hysenlabs.com/projects/jiangtaovan-tdxrs

## The concurrency table shows the pool losing ground as connections are added

The most revealing page in this documentation is the concurrency table, because the numbers contradict what the column heading suggests. It compares two configurations at 5 threads and at 60 threads. The direct client, which opens its own TCP connection per request, goes from 381ms to 344ms and is annotated as showing no degradation. The pooled client goes from 337ms to 4110ms, a factor of 12.2 times. The async client goes from 345ms to 3880ms, a factor of 11.2 times.

So that column is not a speed-up ratio at all; it is how much slower the run becomes when you add threads, and only one of the three keeps that factor below one. Twelve times the threads producing twelve times the wall time means the pooled and async paths bought no throughput whatsoever from the extra connections, while the direct path at least did not go backwards. Neither approaches linear scaling, which for 60 threads at 5 would put the pooled figure in the tens of milliseconds rather than four seconds.

The design conclusion is baked into the client list. The default client is the pooled one with a heartbeat, retries and caching, described as the main path for sequential requests, and a separate direct client exists specifically for the high concurrency case. The comment in the dependency manifest is blunter still, noting that shared connection pools break under load while a fresh connection per call triggers throttling, and prescribing thread-local keep-alive or nothing.

## The rate limiter is loosest when the exchange is shut

Request throttling is built in and described as protecting the server. The ceilings move with the trading session: 15 requests per second during the 9:30 to 15:00 session, 30 per second before and after it, and 60 per second when the market is closed. Two methods control it, one that detects the current session and one that sets it by hand, with the three phases named as trading, prepost and closed.

The logic is deliberate and understandable, since an idle server deserves less protection than a busy one. It has two consequences that are easy to miss. First, any latency measurement you take depends on the wall clock, so a benchmark run at the weekend and one run at ten in the morning are not comparable and neither is reproducible on demand. Second, the limit is enforced per connection, and the documentation says so, noting that four pooled connections multiply the effective throughput by four. A limit designed to protect a server scales with however many connections a client chooses to open, which inverts the stated purpose whenever a caller raises its own connection count.

Batch calls have their own ceiling, with a hard limit of 60 instruments per request and anything beyond that truncated automatically rather than rejected, so a caller passing 200 symbols gets 60 symbols and no error.

## Three parts of the documentation count the same layer three ways

The feature section is headed as network market data covering thirteen categories, and the table beneath it contains eight rows: candlesticks for individual stocks and indices across twelve periods from one minute to yearly, live quotes with five-level order book depth, intraday series both current and historical, tick-by-tick trades with automatic pagination, the full security list with caching, financial data, ex-rights and ex-dividend history, and sector classification.

The architecture diagram counts a third number. Its protocol layer is described as holding eleven parsers plus the adjustment algorithm, and the diagram's own caption refers to five clients and four readers. So the same subsystem is thirteen things in one heading, eight in the table under it, and eleven in the diagram. None of the three is obviously wrong on its own terms, since some of the eight rows expand into several parsers and some of the thirteen may count things the table folded together, but a reader trying to work out what is implemented cannot use any of the counts as a checklist.

The rows that do expand are the informative ones. Financial data is described as thirty-four realtime fields plus forty-five named indicators, which alone would account for a large part of the difference between the table and the diagram.

## The headline example needs an extra the install line does not pull

The very first code block in the documentation connects a client and asks for daily bars as a dataframe. The install section offers two routes and both stop short of what that example needs:

```bash
pip install tdxrs
```

```bash
git clone https://github.com/jiangtaovan/tdxrs && cd tdxrs
pip install maturin
maturin develop --release
```

Those instructions do not fit together with the example. The project metadata declares an empty list of required dependencies, and pandas sits in an optional group described as being for the dataframe-returning methods. So the opening example, which is also the one the description quotes as the reason to use the library, needs an extra that neither install route includes, and the documentation does not mention the extra anywhere in its install steps. A reader following the quick start literally ends up with a missing import rather than a chart.

The rest of the packaging is tidy by comparison. The console script is declared and mapped to a module entry point, so the command line tool arrives with the install, and the build backend is bounded to a major version range. The exclusion list for the published wheel names two files that are not in the repository tree, which costs nothing but suggests the list was not revisited after those files moved.

## Financial values come back raw, with rules of thumb instead of units

The finance interface returns what the server sends, with no unit conversion, and the documentation is honest about it in a comment inside the example: realtime financial values are raw and units are not converted automatically. What follows is a set of empirical rules rather than a specification. Share-capital-like fields are roughly in units of ten thousand, asset-like fields roughly in units of ten thousand, and per-share figures roughly in currency units.

The worked example makes the scale concrete. A net asset value printed as 270894048 is annotated as about 270.9 billion in the larger unit, and a per-share net asset value prints as 216.32. Both arithmetic checks out, which is reassuring, and neither tells you what happens when a field breaks the rule. With forty-five named indicators and thirty-four realtime fields and no unit attached to any of them, a caller has to know which bucket each field falls into before a number means anything.

The same pattern appears in the local file reader, whose constructor takes a scaling coefficient the caller has to supply with the right magnitude, defaulting to none. Nothing in either interface returns the unit alongside the value.

## The engineering summary counts six crates and the manifest lists eight

The documentation closes its highlights with a four-line summary: the language edition, zero lines of unsafe code, a test count, a dependency count, and a documentation count. The dependency line names six core crates.

The manifest names eight. Alongside the six named, the manifest also lists a JSON serialization library and a regular expression engine, and the comment beside the regular expression one says it is there for text parsing of a specific data category. That category has its own feature flag, and the manifest declares three features of which only the default is empty, so two flags control behaviour that the summary does not mention. Development-only dependencies add a temporary file helper and a benchmarking framework configured to emit HTML reports, with a benchmark target declared for the reader.

The release profile is tuned for speed rather than size, with the highest optimisation level and link-time optimisation enabled, and the library is built as both a C-compatible dynamic library for the Python extension and a Rust library so the Rust side can be tested. All of that is consistent with the performance claims; the crate count in the summary is simply out of date.

## The container's default command is an import smoke test

The container definition builds in two stages that are both based on the same slim Python image at version 3.13. The first installs build tooling, then fetches the Rust toolchain by piping a download into a shell interpreter with transport pinned to a modern TLS version, builds the native extension in a virtual environment, and the second copies that environment into a fresh image. The metadata supports two ways to use it: running the image, or building with output written to the current directory to extract the wheel.

The default command of the final image is a one-line Python invocation that imports the module and prints a confirmation. That is a smoke test, not a server and not a tool: the image has no entry point for the command line interface, and no data source is configured. It is a reasonable way to check that a wheel built on your machine imports, and it is worth knowing that is all it does before you assume the image gives you a usable runtime.

The interpreter version is also narrower than the package. The metadata declares support from Python 3.11 upward across three minor versions, so the image tests one of the three it claims to support, and the Windows build carries a separate prerequisite documented elsewhere in the installation notes.

## One README in Chinese, an English file beside it, and two duplicated badges

The documentation this write-up is based on is the Chinese file, and the repository carries a second README in English next to it, with a language switcher at the top of the Chinese one pointing across. Both are in the tree, so a reader wanting the English wording has it, but the two are maintained separately and the page you land on from the repository listing is the Chinese one.

The badge block above the title is worth a second look for a different reason. It holds eight links with no link text, and two of the destinations appear twice: the package index entry is linked at two points in the same block, and the repository address appears both on its own and again with the commits path appended. So a third of that block is redundant.

The rest of the structure is unusually well organised for a project this size. Documentation is split into public and internal sets with separate pages for the benchmarks, the fund module, the command line guide and the installation notes, examples exist in both languages, and a benchmark chart generator sits alongside them.

## Conclusion

Use it if you are parsing large local history files or doing many sequential quotes, since those are the paths the measurements support, and it keeps the tdxpy call shape so existing code moves over with few edits. Do not adopt it for high-concurrency work without reading the concurrency table first, because the pooled client degrades sharply as connections are added and the per-request client is the one that holds up. Three things to check before you rely on it: whether your dataframe calls have the optional extra installed, since the headline example needs it and the one-line install does not pull it; what units the finance fields come back in, since the library returns raw values and only offers rules of thumb; and whether the rate limiter's session clock makes your own benchmarks unrepeatable, since the ceiling is highest when the exchange is shut.

## FAQ

### What is tdxrs?

It is a Rust implementation of the TDX market data parsing library with Python bindings built through PyO3 and maturin, keeping the same API shape as the Python library it replaces. It ships five clients, four local file readers, a downloader and a command line entry point, and its documentation is written in Chinese with an English file beside it.

### How much faster is tdxrs than the Python library it replaces?

The published local file parsing table shows 9 to 11 times, from 0.3ms against 2.8ms for a thousand daily bars up to 0.8ms against 8.5ms for five hundred financial records. Network calls show far less: 1.5 times for a hundred bars and 1.3 times for three quotes. The full numbers are in docs/public/BENCHMARKS.md.

### Which tdxrs client should I use for high concurrency?

The direct client, which opens a separate TCP connection per request. The concurrency table puts it at 344ms across 60 threads against 381ms across 5, while the pooled client goes from 337ms to 4110ms and the async client from 345ms to 3880ms, and the direct client is the one described as showing no degradation at 60 threads.

### Does tdxrs adjust historical prices on the server or on your machine?

On your machine. The server returns unadjusted raw data, and the library applies the standard ex-rights formula locally for forward and backward adjustment, with automatic completion of early ex-rights events, and the unadjusted path is documented as carrying no extra cost.

### What does pip install tdxrs actually pull in?

The project metadata declares no required dependencies at all. Pandas is an optional extra needed by the methods that return a dataframe, so the documented quick start example that builds a dataframe needs an extra the one-line install does not include.

## Sources

- [Issues](https://github.com/jiangtaovan/tdxrs/issues)
- [jiangtaovan/tdxrs on GitHub](https://github.com/jiangtaovan/tdxrs)
- [License: MIT](https://github.com/jiangtaovan/tdxrs/blob/main/LICENSE)
- [README](https://github.com/jiangtaovan/tdxrs/blob/main/README.md)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/jiangtaovan-tdxrs
