DXcam: Desktop Duplication Screen Capture for Windows in Python
A Python high-performance screen capture library for Windows using Desktop Duplication API - Updated 2026
At a glance
- What is it?
- DXcam wraps the Desktop Duplication API and Windows Graphics Capture behind a small Python object, returning numpy arrays at a target frame rate. It is a Windows-only, MIT-licensed library, and the trade-off is that the buffer you get back may not be yours to keep.
- Who is it for?
- DXcam fits Windows-only pipelines that need continuous frames rather than one-off screenshots: video writers, machine learning loops, computer-use agents. It is the wrong tool on macOS or Linux, and a poor fit for scripts that want a cheap still image on a machine where MSS already works.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 9 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on October 1, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What DXcam solves, and who is on the other end of it
Most Python screenshot libraries answer one question: what is on the screen right now. DXcam answers a different one: what is on the screen at 120 or 240 frames per second, without the capture path becoming the bottleneck. The README positions it as a low-latency, high-FPS capture library for Windows, based on the Desktop Duplication API, and explicitly names full-screen Direct3D applications as a supported case.
The audience follows from that. If you are writing a computer-use agent that reads the screen in a loop, a computer vision model that needs a steady stream of frames, or a recorder that writes video from a live desktop, the cost of each capture matters. The README's own framing points at AI agent and computer vision use cases, and the repository topics include computer-vision, deep-learning and low-latency. The library returns numpy arrays, which means the frame drops straight into the rest of the Python numerical stack without a conversion step you have to write.
The project is Windows-only. The classifiers list Windows 10 and Windows 11, and pyproject.toml requires Python 3.10 or newer. There is no macOS or Linux path, and no plan for one is visible in the files.
The capture path: DXGI or WinRT, then a processor backend
DXcam separates capture from post-processing, and the two halves are configured independently.
The capture backend acquires a BGRA frame. The default is dxgi, the Desktop Duplication path; the alternative is winrt, the Windows Graphics Capture path. The README gives a short guideline rather than a benchmark table: start with dxgi for most workloads, especially one-shot grabs, and try winrt if it performs better on your machine or if you need cursor rendering. That is an honest admission that the answer depends on the hardware and the application, and it is also a sign that the project has not settled on a single recommended path.
The processor backend takes that BGRA frame and does rotation and cropping preparation plus color conversion to the output_color you asked for. Supported modes are RGB, RGBA, BGR, BGRA and GRAY. Only BGRA avoids a conversion step, and the README calls it the leanest dependency path because it does not require OpenCV. The other four modes need cv2 or the compiled numpy backend. Binary wheels ship with the Cython kernels those processor backends use, which is why a plain pip install can convert colors without a compiler on the user's machine.
For continuous capture, start() spins up a thread that polls newly rendered frames and stores them in a fixed-size in-memory ring buffer. get_latest_frame() blocks until a frame is available. When capture is running, grab() reads from that ring buffer instead of polling DXGI directly. The buffer defaults to 8 frames and is configurable through max_buffer_len; new frames overwrite old ones when it is full.
Installing DXcam and grabbing a first frame
The minimal install pulls comtypes and numpy, nothing else:
pip install dxcamIf you want OpenCV-based color conversion and the WinRT capture backend, install the extras. Note the quoting, which is what the README shows:
pip install "dxcam[cv2,winrt]"Official Windows wheels are built for CPython 3.10 through 3.14. Building from source is not described in the README; it points at CONTRIBUTING.md.
A first capture is three lines. create() returns a camera bound to the primary output on device 0, and grab() returns a numpy array. Using the context manager releases the capture resources at the end of the block:
import dxcam
with dxcam.create() as camera:
frame = camera.grab()One behaviour to expect on the first run: grab() returns None when no new frame has been rendered since the last capture, which the README describes as backward compatibility. If you want the latest frame regardless, pass new_frame_only=False. For a region instead of the full output, pass a (left, top, right, bottom) tuple to grab(region=...), and the array comes back as (H, W, C).
Before writing capture code, it is worth printing what the machine actually exposes:
import dxcam
print(dxcam.device_info())
print(dxcam.output_info())The output names each device with its VRAM and vendor id, and each output with its resolution, rotation and whether it is primary. Those are the indices you pass as device_idx and output_idx when creating cameras.
The ring buffer, zero-copy views and the trap they set
The performance story depends on not copying pixels, and DXcam exposes that directly. grab(copy=False) and grab_view() return a zero-copy view into the frame buffer. The README is explicit about the consequence: the returned buffer can be overwritten by later captures. If you hold that array while the capture thread keeps running, you are reading memory that is being rewritten underneath you. For a model that consumes a frame and finishes before the next one arrives, this is fine. For anything that queues frames, compares two frames, or stores them, it is a bug waiting to happen, and it will not announce itself.
The ring buffer has the same character. With max_buffer_len=8 (the default), a consumer slower than the producer silently loses frames rather than applying backpressure. That is the correct design for live capture, where the newest frame is the only one that matters, and the wrong design for anything that needs every frame. The README does not document a way to detect dropped frames, and the timestamp returned by get_latest_frame(with_timestamp=True) is the frame's presentation time, not a sequence number, so gaps are not directly visible from the API surface described.
video_mode=True changes the contract: the buffer is filled at the target frame rate, reusing the previous frame when nothing new was rendered. That is what you want when writing a video file at a fixed frame rate, and it is misleading if you are using frames as observations of a changing screen, because a repeated frame looks identical to a static screen. The README's own video example uses video_mode=True with a cv2.VideoWriter at 30 fps, which is the intended pairing.
Where DXcam is the wrong tool
The clearest limitation is the platform. DXcam is Windows only, full stop. On Linux or macOS the question is not whether DXcam is fast, it is that there is no DXcam.
The second limitation is the beta status. pyproject.toml carries the classifier Development Status :: 4 - Beta. The release history is short and recent: v0.1.0 on 2026-03-08, v0.2.0 on 2026-03-10, v0.3.0 on 2026-03-12. Three releases inside a week, followed by no release for months. The last push to the repository was on 2026-03-18. That is not an abandoned project, but it is also not a project with a long track record of API stability, and the README already documents one compatibility decision (grab() returning None) that exists to avoid breaking callers.
The third is resource cost. The README itself warns that target_fps greater than 120 is resource heavy. Continuous capture at high frame rates occupies a thread, holds GPU-side duplication resources, and consumes CPU in the processor backend for any output_color other than BGRA. On a laptop on battery, or on a machine that is also running the workload you are capturing, this is a real cost.
Finally, release() is terminal. After release(), the same instance cannot be reused and start() raises RuntimeError. Code that tries to restart a camera after releasing it will fail, and the README does not document a way to revive the object.
DXcam against MSS and BetterCam
The comparison people actually search for is DXcam versus MSS, and the difference is architectural rather than a matter of tuning.
MSS reads the screen through the platform's GDI or equivalent path and returns a still image on demand. It is cross-platform, it has no capture thread, no ring buffer and no notion of a target frame rate, and every call is a fresh read you own. That model is simple and predictable, and for a script that takes a screenshot every few seconds it is entirely sufficient.
DXcam goes through Desktop Duplication, which is the same mechanism the desktop compositor uses to hand frames to consumers. It gives you a stream, timestamps derived from DXGI_OUTDUPL_FRAME_INFO.LastPresentTime on the dxgi backend or from WinRT SystemRelativeTime on the winrt backend, and pacing with drift correction to hold near target_fps. The cost is that you now manage a lifecycle: create, start, consume, stop, release. If your workload is one screenshot per minute, that lifecycle is overhead with no payoff, and MSS is the better choice.
BetterCam appears in the related searches as a sibling project in the same space, and the README does not mention it, so no comparison can be made from this material. The same applies to wincam and Dxcam-CPP. What can be said is that DXcam's distinguishing feature within its own documentation is the dual backend, dxgi or winrt, which gives you a second capture path to try when the first one misbehaves on a particular machine.
Maintenance, licence and what an upgrade costs
The repository is not archived. The last push was on 2026-03-18, and the newest tagged release is v0.3.0 from 2026-03-12. A CHANGELOG.md exists at the top level, and it is the file to read before moving between versions, because the README does not summarise what changed in each release.
Upgrade cost is mostly about the Python version and the extras. requires-python is >=3.10, and wheels are built for 3.10 through 3.14. Moving to a Python version outside that window means building from source, and the README only points at CONTRIBUTING.md for that. The cv2 and winrt extras are separate dependency sets: winrt pulls a list of winrt-Windows.* packages at >=3.2.1, and cv2 pulls opencv-python. If you installed the minimal package and later switch output_color from BGRA to BGR, you need the cv2 extra or the compiled numpy kernel, and that is a dependency change rather than a code change.
The licence is MIT, declared both in pyproject.toml and in the LICENSE file. MIT is permissive: it allows commercial and closed-source use, and it requires that the copyright notice and permission notice be included in copies or substantial portions of the software. It provides no patent grant and no warranty. That is a description of the licence text, not legal advice; if the capture path is part of a shipped product, have someone qualified read the actual LICENSE file.
Editorial conclusion
DXcam fits Windows-only pipelines that need continuous frames rather than one-off screenshots: video writers, machine learning loops, computer-use agents. It is the wrong tool on macOS or Linux, and a poor fit for scripts that want a cheap still image on a machine where MSS already works. Before adopting, check dxcam.device_info() and dxcam.output_info() on the target machine, confirm which backend and output_color your hardware tolerates, and read the release notes for v0.3.0 to see what changed from v0.2.0.
Frequently asked questions
What is the best Python library for screen capture?
There is no single answer, and DXcam's own README frames it as a choice rather than a verdict: DXcam targets low-latency, high-FPS capture on Windows through Desktop Duplication, while the README's comparison section positions it against common Python alternatives on throughput, full-screen Direct3D stability and FPS pacing. If your workload is a still screenshot every few seconds on a cross-platform machine, DXcam's capture thread and ring buffer are overhead you do not need.
How do I install DXcam with pip?
The minimal install is pip install dxcam, which brings in comtypes and numpy. For OpenCV-based color conversion and the WinRT capture backend, the README shows pip install "dxcam[cv2,winrt]". Official Windows wheels are built for CPython 3.10 to 3.14.
Why does DXcam grab() return None?
grab() returns None when no new frame has been rendered since the last capture, which the README describes as behaviour kept for backward compatibility. Passing new_frame_only=False makes DXcam always return the latest frame instead.
Does DXcam work on macOS or Linux?
No. DXcam is built on the Desktop Duplication API and Windows Graphics Capture, and pyproject.toml classifies it for Windows 10 and Windows 11 only. The README documents no other platform.
What is the difference between the dxgi and winrt backends in DXcam?
dxgi is the default Desktop Duplication path and the README recommends starting there, especially for one-shot grabs. winrt is the Windows Graphics Capture path, and the README says to use it if you need cursor rendering or if it performs better on your machine.
What does DXcam's release() do?
release() stops capture, frees buffers and releases capture resources. After release(), the same instance cannot be reused and calling start() raises RuntimeError. Creating the camera with a with block releases it automatically at the end of the block.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/ra1nty-dxcam)