# kcp-go: reliable UDP for Go, tuned down to the nanoseconds

> kcp-go is xtaci's MIT-licensed Reliable-UDP library for Go with FEC support, providing smooth, resilient, ordered, error-checked and anonymous stream delivery over UDP, compatible with net.Conn as a drop-in replacement for net.TCPConn. Deployed across millions of devices from MIPS routers to servers, it handles over 5K concurrent connections per commodity server, encrypts every packet with AES, Salsa20 or AEAD, and documents its design decisions with benchmarks.

**xtaci/kcp-go** — A crypto-secure Reliable-UDP library for Golang with FEC support.

- Repository: https://github.com/xtaci/kcp-go
- Stars: 4,557 · Forks: 809
- Language: Go
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/xtaci-kcp-go

## The protocol over UDP, and the Go module that implements it

KCP is a reliable protocol that runs over UDP, trading some bandwidth for reduced latency compared with TCP, and kcp-go is its Go implementation, a Reliable-UDP library providing smooth, resilient, ordered, error-checked and anonymous stream delivery over UDP packets. The library is compatible with the C implementation by skywind3000 with various improvements, and its deployment claim spans millions of devices from low-end MIPS routers to high-end servers, across online games, live broadcasting, file synchronization and network acceleration. The MIT-licensed package is versioned as v5 on the Go module path, and the badge table links the KCP-Powered mark back to the original protocol project, the lineage stated rather than implied. The complete API documentation lives in Godoc for the v5 module path, and the echo example in the examples directory is the runnable reference for the session layer's client-server shape.

## The feature list, from FEC to sendmmsg

The ten features read as a systems checklist. Designed for latency-sensitive scenarios, cache-friendly and memory-optimized, handling over 5K concurrent connections on a single commodity server. Compatible with net.Conn and net.Listener, serving as a drop-in replacement for net.TCPConn, the API decision that makes adoption mechanical. FEC support using Reed-Solomon codes, recovering lost packets without retransmission round trips. Packet-level encryption for AES, TEA, 3DES, Blowfish, Cast5 and Salsa20 in CFB mode plus AEAD support, generating completely anonymous packets. Only a fixed number of goroutines for the entire server with context switching costs considered. And platform-specific optimizations, sendmmsg and recvmmsg on Linux, reflected in the repository's tx_linux.go and readloop_linux.go files beside their generic counterparts.

## Slice versus list: a 43 percent CPU difference

The key design considerations section benchmarks its own decisions, and the first is the send queue structure. The flush routine loops through the send queue for retransmission checking every 20 milliseconds, and the author's benchmark compares a slice loop against a container/list loop:

```
BenchmarkLoopSlice-4   	2000000000  	         0.39 ns/op
BenchmarkLoopList-4    	100000000   	        54.6 ns/op
```

0.39 nanoseconds per operation against 54.6. The list structure introduces heavy cache misses compared with the slice's locality, and the arithmetic follows, for 5000 connections with a 32-window size at a 20 millisecond interval, a slice costs 6 microseconds, 0.03 percent of CPU, per flush, while a list costs 8.7 milliseconds, 43.5 percent. A data structure choice measured in orders of magnitude, published with the numbers, is the documentation style that separates a tuned library from an assembled one.

## Timing accuracy, counted in cycles

The second consideration addresses the RTT estimator, since inaccurate timing leads to false retransmissions in KCP, but time.Now costs 42 cycles, 10.5 nanoseconds on a 4 GHz CPU and 15.6 on the author's MacBook. The benchmark for the call is published, and the design minimizes its impact, after each output call the current clock time is updated on return, and a single flush queries the system clock once. The arithmetic quantifies both directions, 5000 connections cost 78 microseconds of fixed clock reading when no packets are pending, and a 10 MB per second transfer at 1400 MTU invokes output roughly 7500 times per second, costing 117 microseconds of time.Now per second. Every clock read that matters is accounted for, the RTT estimator's inputs protected from jitter the library itself would introduce.

## One buffer pool to rule the queues

Memory management starts from a global buffer pool, xmit.Buf, and allocations return fixed-capacity 1500 byte blocks, the mtuLimit. The rx queue, tx queue and FEC queue all receive bytes from this pool and return them after use, preventing unnecessary zeroing, and the pool maintains a high watermark for slice objects so in-flight objects survive periodic garbage collection while memory can return to the runtime when idle. The design eliminates per-packet allocation from the hot path, the difference between a library that measures well in microbenchmarks and one that holds up under thousands of connections, and the ringbuffer.go and bufferpool.go files in the repository carry the implementation with their own tests beside them.

## Anonymity down to the headers

The security design is stated as information security, with built-in packet encryption across the block ciphers in Cipher Feedback mode, and the construction matters, for each packet the encryption process begins by encrypting a nonce from system entropy, ensuring the same plaintext never produces the same ciphertext. The stronger claim follows, packet contents are completely anonymous with encryption, including the headers, FEC and KCP, the checksums and the payload, so an observer cannot even identify the protocol from the wire. The entropy module in the repository backs the nonce source, the gmsm dependency adds Chinese national standard cryptography, and the AEAD path provides authenticated encryption for deployments that need integrity with their confidentiality, the three levels of the security story, none, CFB anonymous, AEAD.

## Fixed goroutines, timed scheduling, and autotune

The design considerations continue into packet clocking and FEC characteristics, the sections covering how the protocol's timing loops are structured, and the feature list states the concurrency model plainly, only a fixed number of goroutines are created for the entire server application, with context switching costs between goroutines taken into consideration. The repository's file list names the machinery, readloop and tx split per platform, timedsched for the timing scheduler, autotune for parameter adaptation, and a wireshark directory providing protocol dissection for debugging on the wire. The sess.go file implements the session layer presenting net.Conn, and the echo example in examples demonstrates the client-server shape in minimal code, the entry point the Godoc complements.

## Maintenance, SNMP metrics, and a 2026 release

The library's maintenance surface is visible in its details, an snmp.go exposing protocol statistics, coverage reporting through codecov, a Go Report Card badge and Sourcegraph integration, and the go.mod targeting Go 1.24 with dependencies kept to klauspost's reedsolomon for the FEC math, golang.org/x/crypto, x/net, x/sys and x/time, plus the lossyconn test package for simulating lossy networks. The latest release tag, v5.6.64 in January 2026, reflects a versioning scheme that has accumulated through years of point releases, with the master branch pushed 2026-05-15. An AGENTS.md file configures coding agents for contribution, and the FAQ and who-is-using sections of the documentation round out the operational knowledge a networking library needs to be trusted.

## Conclusion

Use kcp-go when latency-sensitive traffic, game state, live media, synchronization, must survive lossy networks where TCP's retransmission behavior hurts, and when the Go net.Conn interface should be preserved so existing code swaps transports. It is the wrong tool when plain UDP or QUIC already fits, or when the application is bandwidth-bound rather than latency-bound. Before adopting, read the key design considerations for the performance contract, choose the encryption mode deliberately since packet anonymity includes headers, plan for the fixed number of goroutines rather than per-connection ones, and check the specification section against your reliability requirements.

## FAQ

### What is the KCP protocol?

KCP is a reliable protocol that runs over UDP, trading bandwidth for lower latency than TCP, and kcp-go is its Go implementation. The library delivers smooth, resilient, ordered, error-checked and anonymous streams over UDP, with Reed-Solomon FEC and packet-level encryption, compatible with Go's net.Conn interface.

### What does KCP stand for?

KCP is the name of the reliable UDP protocol by skywind3000 that kcp-go implements in Go, compatible with the original C version with various improvements. The library positions itself as KCP-Powered, carrying the protocol's focus on latency-sensitive scenarios like online games, live broadcasting and network acceleration.

### How many connections can kcp-go handle?

The library states it handles more than 5,000 concurrent connections on a single commodity server, with only a fixed number of goroutines created for the entire server application and context switching costs taken into consideration, supported by the buffer pool and slice-based send queue design.

## Sources

- [Issues](https://github.com/xtaci/kcp-go/issues)
- [License: MIT](https://github.com/xtaci/kcp-go/blob/master/LICENSE)
- [README](https://github.com/xtaci/kcp-go/blob/master/README.md)
- [Releases](https://github.com/xtaci/kcp-go/releases)
- [xtaci/kcp-go on GitHub](https://github.com/xtaci/kcp-go)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/xtaci-kcp-go
