grok2api-egress-enhancements, three delivery paths for one feature set
Grok2API & CPA egress quality guard with proxy recovery, quarantine, migration, and operations UI
At a glance
- What is it?
- grok2api-egress-enhancements is an unofficial distribution layer over chenyme/grok2api that adds egress quality guardrails: hold a streaming response until real reasoning appears, quarantine proxy nodes that fail a quality check, and recover fixed proxies quickly. It ships the same behaviour three different ways depending on which upstream version you are on, and its own documentation is written as a prompt to paste into an AI assistant.
- Who is it for?
- This repository is for operators already running chenyme/grok2api across many accounts and many residential exits, where a silently degraded model is a billing problem rather than a quality preference. It is not a general client.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 36 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 2, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Three delivery paths and a warning about mixing them
One feature set, three ways to install it, and a specific warning about applying the wrong series.
If you are on current upstream, the README says the relevant patches are already merged. The runnable tree is the fork lij768423-svg/grok2api, cloned and built rather than pulled.
If you are still on a clean v3.1.4, you apply six patches from the patches/ directory:
git fetch --tags origin
git checkout -b tui-hold-serde v3.1.4
git am --3way /path/to/grok2api-egress-enhancements/patches/0015-fix-quality-thinking-evidence.patch
git am --3way /path/to/grok2api-egress-enhancements/patches/0016-fix-responses-annotations.patch
git am --3way /path/to/grok2api-egress-enhancements/patches/0017-fix-hold-tui-tool-turns.patch
git am --3way /path/to/grok2api-egress-enhancements/patches/0018-fix-empty-completed-retry.patch
git am --3way /path/to/grok2api-egress-enhancements/patches/0019-fix-clear-account-cooldown.patchIf you are on v3.0.11, a single legacy patch applies, the one associated with the pull request that is now closed.
The warning is specific: do not apply patches 0006 through 0014 on top of v3.1.4. Those belong to a different baseline. Yet the release title for v1.1.0 in this repository is described as grok2api lab patches 0006 to 0009, so the repository's own version history and its patch numbering are two separate series, and reading one as a guide to the other is an easy mistake.
The same parameters, enabled by default here and off upstream
The interception logic matches the upstream pull request exactly. Only the default differs.
The README is explicit that the parameters are the same as the ones the official pull request recommends: a hold timeout of 30 seconds, a minimum visible output of 8 tokens, a cipher floor of 256 bytes, and a multiplier of 4 per reasoning token. What differs is a single switch. requestRetry.enabled defaults to false upstream and to true in the fork, and the fork is described as enabling the interception on by default.
That matters for anyone evaluating the patch. Applying it upstream-style gives you the code with the behaviour dormant, and the default is what a stock deployment will use.
The hold itself is narrower than it sounds. A streaming request from a reasoning model is held before anything is written to the user, and only real streaming thinking releases it: a reasoning or summary delta, or encrypted_content or an Anthropic signature that has reached the cipher floor. A short stub does not count. A placeholder-looking encrypted value, an empty reasoning item, a stub event, or a reasoning token count appearing on its own are all explicitly not evidence.
The floor is computed as the greater of the minimum encrypted bytes and the reasoning token count times the per-token multiplier, which defaults to 256 bytes and 4.
A flush just over the floor still gets held
The edge cases are documented in more detail than the feature itself, which is where the engineering effort went.
Three specific cases extend the hold. First, if the hold expires and the response is still empty, a short greeting that arrives later with a high reasoning count continues to be held; the expiry state is written onto scan state specifically so it is not lost on the timer tick. Second, a flush that arrives within one second of clearing the floor is held, which is the case where a provider trickles the minimum and then emits the real answer immediately. Third, TUI turns keep holding: the Grok TUI sends a tools schema on every turn, and once a tool has run the next turn already carries function_call_output, so both of those shapes stay held because switching account only replays the model turn.
Hosted tools are treated differently and are not replayed. That distinction is the difference between a hold that can retry safely and one that cannot.
An empty completion event switches account immediately rather than waiting for an idle timeout before returning HTTP 200. At most six shots are tried, and if none produced reasoning the request returns 503 with a quality_degraded code. The first miss puts the account on a twelve hour cooldown; a second miss after that cooldown disables it outright. An empty stream is penalised separately with a fifteen minute idle cooldown. TUI compression is exempt from the hold entirely.
The quality guard calls itself a heuristic, not an adjudicator
The documentation contains its own strongest caveat, and it is worth reading before tuning anything.
The quality guard is described as a heuristic circuit breaker rather than a model capability adjudicator. Middle-layer buffering, an existing file, a long constant, or cached content can all produce an abnormally high instantaneous token rate. The hard threshold policy is called aggressive, with the advice to raise it per link, while the soft threshold still confirms with a fixed prompt retest.
That is an honest framing of a measurement problem. The metric is output tokens divided by the total duration minus the time to first token, with reasoning tokens counted in the output, which is exactly the number that a buffering proxy inflates.
The thresholds are named: soft at 500 and hard at 1000 tokens per second. A generation window shorter than one second that reaches the soft threshold is recorded under a separate category, buffered_burst, so the two are distinguishable in the panel rather than merged.
Two behaviours follow from the heuristic framing. A hard hit quarantines the node immediately. A soft hit triggers an active retest with a fixed prompt, and the node is quarantined only after consecutive hits, which is the confirmation step that keeps buffering from causing an outage.
A new IP is tested exactly once before it is trusted
Node replacement is where most of the operational logic sits, and it is conservative in a specific way.
A replacement IP gets one real model quality check and no more. If it passes, the node recovers immediately. If it fails, or if the result is uncertain, the node stays quarantined. There is no second attempt and no optimistic restore.
The rotation machinery is named down to the session format. A trusted node-level webhook for changing IP is supported, and there is a specific sticky session rotator for a provider whose session identifiers take the shape sid-...-t-....
Strict mode drains traffic before confirming, and a short-window streaming burst that could be a false positive is retested on the same IP first, only switching after the anomaly is confirmed. That ordering matters: without it, a single buffering artefact would cost you a good IP.
Scheduling failures and proxy failures are kept apart on purpose. When no account can be scheduled, the retest is deferred rather than counted as a proxy error and rather than spending traffic on a new IP. When the whole pool is unavailable, detection is deferred on a separate long backoff with duplicate log suppression, and the node stays quarantined while it waits. Where a probe needs an account, it will borrow any healthy one, while the measured request is still forced through the node under test.
Probe profiles hide their prompt text from the state API
The probe configuration has a small detail that is easy to miss and easy to appreciate.
A quality guard page carries probe profiles, with a built-in expected marker consisting of a last line of QUALITY_OK and a throughput baseline. Custom profiles can be built from a prompt, a contains match, a last line, or a regular expression.
The behaviour differs by what the probe finds. A missing marker is a hard anomaly. A short reply that does contain the marker is not quarantined for inflated throughput or for too few tokens, which is the guard against the false positive the previous section described.
The profiles live in a file called profiles.json alongside the runtime configuration. The state API returns profile names and whether a profile has a marker, and deliberately does not return the prompt or the marker body. So an admin page can show you which profiles exist without shipping their contents to every browser that renders it.
The interface is four verbs on one path: GET and POST on /api/admin/v1/egress-quality-guard/profiles, and PUT and DELETE on a single profile by identifier. A quality check can carry a profile identifier. The degraded accounts panel has its own read-only endpoint for classifying accounts from request audit data.
The install instructions are a prompt for an AI assistant
The first section of the documentation is a block of text to paste into a chat window, and that is a deliberate choice.
The file calls it a one-click install prompt: copy the whole segment, send it to your AI, and change only the residential addresses at the end. The same block is said to exist in the fork's README.
The prompt itself is unusually opinionated for documentation. Do not improvise. Do not install the CPA plugin. Do not pull the published upstream image, because the fork has to be cloned and built with docker compose up -d --build. Use every residential address you have, one Mihomo listener and one Grok2API node per sticky session, and it explicitly forbids enabling a single node and forbids merging several into one residential pool.
The reasoning behind that is given in one line above the prompt: to approximate how this runs in production you hand over the full set of home connections, not one node.
The rest of the repository is conventional. There is a patches/ directory, a scripts/ directory with a Python helper that splits residential addresses into listeners and nodes, a standalone sidecar/quality_guard.py paired with the admin page at /quality-guard, and docs/ with an installation guide and a recommended deployment chain. Documentation is duplicated in Chinese and English, and SECURITY.md sits next to them. The repository homepage is a pull request URL, and that pull request is closed.
Editorial conclusion
This repository is for operators already running chenyme/grok2api across many accounts and many residential exits, where a silently degraded model is a billing problem rather than a quality preference. It is not a general client. Check four things before adopting it. Which upstream version you are on, because the delivery path differs at v3.0.11, v3.1.4, and current, and the README warns explicitly against applying the wrong patch series. Whether you have the hardware for the fork build rather than the published image, since pulling the upstream image gives you interception disabled. Whether your hard token-per-second threshold is set too low for your link, because the project says the heuristic is aggressive and buffering can inflate the number. And whether your admin pages are reachable only where the sidecar is running.
Frequently asked questions
What does grok2api-egress-enhancements actually add?
It is an unofficial enhancement distribution for chenyme/grok2api that does not copy upstream source. It adds fixed proxy fast recovery, an egress quality guard with node quarantine and optional sticky IP rotation, a Python sidecar paired with the admin page at /quality-guard, and an interception that holds streaming responses until real reasoning appears.
How do I install grok2api-egress-enhancements?
Three ways depending on version. On current upstream the patches are already merged, so clone the lij768423-svg/grok2api fork and build it. On a clean v3.1.4 apply patches 0015 through 0020 with git am. On v3.0.11 use the single legacy patch. The README warns against applying patches 0006 to 0014 on top of v3.1.4.
Why does grok2api-egress-enhancements quarantine a proxy node?
Passive audit computes output tokens over total duration minus time to first token, with reasoning tokens included, against thresholds that default to 500 and 1000 tokens per second. A hard hit quarantines immediately; a soft hit triggers an active retest with a fixed prompt and quarantines only after consecutive hits.
Is the CPA plugin part of grok2api-egress-enhancements?
No, it is not the default deliverable. It lives in cpa-plugin/ and is a pure native CPA plugin that does not depend on or connect to the Grok2API runtime. The documentation states a single-account or stable-static-proxy setup can skip installing it.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/lij768423-svg-grok2api-egress-enhancements)