Library / SDK
dotnet/Open-XML-SDK avatar
dotnet/Open-XML-SDK

dotnet/Open-XML-SDK: the layer between C# and the inside of a docx file

Open XML SDK by Microsoft

4,613 stars608 forksC#MIT

At a glance

What is it?
Microsoft's open source library for reading and writing Word, Excel and PowerPoint files at the package and markup level, split across four NuGet packages and shipped on a daily build feed.
Who is it for?
What decides whether this SDK is a good fit is not whether you like C# but whether you understand the file format. The README states the prerequisite plainly and links to the Microsoft Office implementation notes for ISO 29500, and that sentence is the real cost of the library.
Can I use it commercially?
Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
Is it still maintained?
Yes. The repository last received commits 1 day ago.
What is it written in?
Mainly C#, according to GitHub's language statistics.

Answers come from the project's GitHub data, last synced on October 7, 2026, and from our analysis. They are not legal advice.

Editorial analysis

A format library wearing an SDK badge

The README is unusually direct about what this repository is and is not. It calls the Open XML SDK a framework for working with Microsoft Office Word, Excel and PowerPoint documents, and then narrows that to APIs for low level operations related to OPC packages, Flat OPC files, and Open XML markup in two forms, strongly typed classes and LINQ to XML. That is the entire product surface described in one sentence.

What follows immediately is the boundary. The SDK is designed to closely follow the Microsoft Office implementation of the ISO 29500 standard, but it was not intended to directly provide higher level abstractions or productivity tools. Read that twice if you are evaluating the library. There is no document builder that reads like English, no template engine, no spreadsheet formula parser standing on its own. The abstraction offered is the format itself, exposed as types.

The repository is under the dotnet organization, written in C#, licensed MIT, with 4,607 stars and 606 forks at capture time. It is not archived, `main` is the default branch, and the last recorded push is 2026-09-16. Roughly 130 open issues sit against that, which is a normal ratio for a library this central to enterprise document handling. Topics listed are docx, office, openxml-format and pptx. The homepage points at the DocumentFormat.OpenXml package on NuGet, which is a hint that the package, not the repository, is the thing most people consume.

The five scenarios the README promises

The README lists what the APIs enable, and the list is short enough to be worth quoting in full because it tells you exactly which workloads the library was designed for.

code
- High-performance generation of word-processing documents, spreadsheets, and presentations.
- Document modification, such as adding, updating, and removing content and metadata.
- Searching and replacing content using regular expressions.
- Splitting up (shredding) a file into multiple files, and combining multiple files into a single file.
- Updating cached data and embedded spreadsheets for charts in Word/PowerPoint.

Three of those five are about bulk work rather than editing a single document. High performance generation is the claim that matters most for server side work, where a report run producing hundreds of files needs to stay predictable in memory. Shredding and combining is the scenario that shows up in mail merge style systems and in archive tooling. Updating cached chart data is the one that catches people out, because a chart embedded in a Word document carries its own copy of the underlying spreadsheet, and keeping that copy in sync is a separate operation from editing the chart.

Regular expression search and replace is listed as a first class capability rather than an afterthought, which tells you the markup is meant to be walked directly. That is only true because of the second API form the README mentions, LINQ to XML, which sits in its own package.

Four packages instead of one

The packages table in the README lists four NuGet packages, each with a stable and a prerelease badge. They are DocumentFormat.OpenXml.Framework, DocumentFormat.OpenXml, DocumentFormat.OpenXml.Linq and DocumentFormat.OpenXml.Features.

The split matters for build size and for dependency hygiene. DocumentFormat.OpenXml.Framework carries the shared plumbing, the validation layer, and the package handling that the other packages build on. DocumentFormat.OpenXml is the generated strongly typed object model, and given the tree contains generated/ and gen/ directories alongside data/, most of its size comes from schema generated code covering the Office markup vocabulary. DocumentFormat.OpenXml.Linq is optional and only needed if you want to query document content with LINQ. DocumentFormat.OpenXml.Features is the newest of the four by name and is where feature level behavior lives.

A plain consumer who only opens and edits a spreadsheet needs the first two. Adding LINQ queries over paragraphs or cells pulls in the third. That packaging decision is the clearest signal about how the maintainers think about this library: it is infrastructure, and infrastructure gets decomposed.

Daily builds on an Azure blob feed

Stable releases go to NuGet.org. Everything else goes to a custom feed hosted on Azure blob storage, and the README includes the exact NuGet.config fragment needed to consume it.

xml
<?xml version="1.0" encoding="utf-8"?>
<configuration>
  <packageSources>
    <add key="OpenXmlCI" value="https://ooxml.blob.core.windows.net/feed/index.json" />
  </packageSources>
</configuration>

Two things to note. First, the feed is an Azure blob with an index.json service index rather than a plain directory listing, so it behaves like a real NuGet v3 source. Second, the README carries an important callout saying the CI feed URL changed as of 2 April 2024, and anyone still pointing at the old address needs to update it. Stable packages on this feed are mirrored onto NuGet and are identical to what you get from NuGet.org, so the feed is purely for prerelease and in flight work.

The practical use is testing against a fix before it ships. If a schema bug or a serialization defect is blocking you, the daily feed lets you take the newest build, verify your scenario, and decide whether to wait for the next tag. CHANGELOG.md is the companion file for tracking what landed.

Two known issues that shape your architecture

The README names both known issues explicitly, and both are architectural rather than cosmetic.

The first is streaming. On .NET Core, .NET 5 and following, ZIP packages have no way to stream data, so the working set can explode in certain situations. The README links this to the .NET runtime issue tracker. If your service reads a 300 MB PowerPoint deck into memory to change one slide, the SDK will not stop you, and the fix is not a flag but a change in how you reason about document size. Either work on a smaller input, or move the processing out of process.

The second is an IsolatedStorageException on .NET Framework. It shows up under certain circumstances when manipulating a large document in an environment with an AppDomain that does not have enough evidence. The mitigation is already in the tree: there is a samples/IsolatedStorageExceptionWorkaround directory, so the repository ships a working example rather than only documenting the problem.

Neither issue is a reason to avoid the SDK. Both are reasons to know which target framework you are on before you size a workload, because the failure modes differ between .NET Framework and modern .NET.

What the recent releases reveal about direction

The three most recent tags tell a coherent story about where effort is going: size, speed, and schema freshness.

Version 3.5.1 from March 2026 is small and purely additive. It added DocumentFormat.OpenXml.Office2016.Drawing.ChartDrawing.Offset, plus Version, FeatureList and FalbackImg attributes on ChartSpace and an ExtensionDropMode enum. That is schema catching up with newer Office markup, nothing more.

Version 3.4.1 from January 2026 is the interesting one. It added MediaDataPartType.Mp4 so MP4 video media parts work in presentations and documents, and updated bundled schemas to the Q3 2025 Office release. The performance work is more significant: the generic builder pattern was removed from element metadata creation to reduce JIT and AOT size and improve document load performance, and FromChunkedBase64String was optimized for roughly 2.4 times the throughput with about 70 percent less memory. Correctness fixes include switching to XmlDOMTextWriter instead of XmlWriter.Create for certain XML content, a clear error when opening encrypted documents the SDK cannot process, and exception messages that name the missing package part instead of failing vaguely.

Version 3.3.0 from March 2025 added the WorkbookCompatibilityVersion namespace and greatly improved ToFlatOpc performance for large parts. The pattern across all three is a library that has moved past feature parity and into making itself viable for large scale server workloads.

The ecosystem that grew up around it

The related tools section of the README is the fastest way to understand the shape of the community, and the entry point tells you something important: the 2.5 Productivity Tool is still listed, with the caveat that it generates code compatible with SDK 2.5. That tool predates version 3 by several major versions and its existence tells you how much churn the object model has seen.

The current generation of libraries all take a different approach. ClosedXml provides a simplified object model on top of the Open XML SDK for Excel. OfficeIMO does the same for Word, Excel and PowerPoint and adds Markdown, Html, Pdf and Visio support. OpenXML-Office covers presentations and spreadsheets. Open XML Powertools provides example code and guidance rather than a wrapped API.

There is also a tool pointing the other direction: Serialize.OpenXml.CodeGen converts an existing Open XML document into the .NET code required to create it. That is a reasonable escape hatch when you want a static starting point for a template instead of code that walks the object graph at runtime.

The samples directory in the tree shows the same practical bent, with Linq, RichData, SVGExample, SunburstChartExample, ThreadedCommentExample, NamedSheetView, AnimatedModel3DExample, DocumentTaskExample and the IsolatedStorageExceptionWorkaround already in place.

Editorial conclusion

What decides whether this SDK is a good fit is not whether you like C# but whether you understand the file format. The README states the prerequisite plainly and links to the Microsoft Office implementation notes for ISO 29500, and that sentence is the real cost of the library. If you are comfortable reading a spec, you get typed access to every part of a document, fine grained control over what gets written, and the ability to shred or recombine files that no spreadsheet library offers. If you want to set a cell value and move on, ClosedXml or OfficeIMO sit on top of this same SDK and save you the detour. The known issues are worth planning around rather than discovering: ZIP packages cannot stream on .NET Core and .NET 5 or later, so the working set can grow unexpectedly, and .NET Framework can throw an IsolatedStorageException on large documents in constrained AppDomains. The version 3.5.1 release is small and additive, while version 3.4.1 is where the engineering shows, with roughly 2.4 times faster chunked base64 decoding, about 70 percent less memory in that path, and a move away from the generic builder pattern to cut JIT and AOT size. Start with the stable DocumentFormat.OpenXml package from NuGet, add the framework and Linq packages when you need them, and reserve the Azure blob feed for testing a change before it reaches a stable release.

Frequently asked questions

What is Microsoft Office Open XML?

Open XML is the file format behind .docx, .xlsx and .pptx files, defined by the ISO 29500 standard. A Word document is not a single flat document but an OPC package, a ZIP container of XML parts with defined relationships between them, which is why a library that understands packages and parts is needed to work with these files programmatically.

How can I use Open XML in C#?

Add the `DocumentFormat.OpenXml` NuGet package, or `DocumentFormat.OpenXml.Framework` alongside it, then open the file with `WordprocessingDocument.Open` or `SpreadsheetDocument.Open` and walk the typed object model. Add `DocumentFormat.OpenXml.Linq` if you want to query content with LINQ to XML. The README asks for detailed knowledge of ISO 29500 as a prerequisite, since the object model mirrors the format rather than hiding it.

Which .NET versions can the Open XML SDK stream documents on?

None of the modern ones. The README lists as a known issue that on .NET Core, .NET 5 and later, ZIP packages have no way to stream data, so the working set can explode in certain situations. On .NET Framework the separate known issue is an `IsolatedStorageException` on large documents in an AppDomain without enough evidence, for which the repository ships a workaround sample.

Official sources

  1. dotnet/Open-XML-SDK on GitHub
  2. License: MIT
  3. Project website
  4. README
  5. Releases
Add this badge to your README

If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.

Add this badge to your README

markdown
[![Hysen Labs](https://hysenlabs.com/badge/dotnet-open-xml-sdk.svg)](https://hysenlabs.com/projects/dotnet-open-xml-sdk)