.NET for Apache Spark: C# and F# Bindings for Spark DataFrames and Structured Streaming
.NET for Apache® Spark™ makes Apache Spark™ easily accessible to .NET developers.
At a glance
- What is it?
- Microsoft's .NET for Apache Spark exposes SparkSQL, DataFrame and Structured Streaming APIs to C# and F# through a .NET Standard library, with a JVM bridge underneath. It is useful if your team already writes .NET and needs Spark; it is not a Spark replacement, and it is not for you if you need Spark 4 or a pure-JVM stack.
- Who is it for?
- Adopt .NET for Apache Spark if your team writes C# or F# and the data platform is already Apache Spark 2.4 through 3.5, particularly on HDInsight, EMR or Databricks where the deployment docs cover the setup. Do not adopt it if you need Spark 4, if you need MLlib, GraphX or RDD-level APIs from .NET, or if your team is comfortable in Scala or PySpark, since those get new Spark features first.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 2 days ago.
- What is it written in?
- Mainly C#, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 29, 2026, and from our analysis. They are not legal advice.
Editorial analysis
The gap .NET for Apache Spark fills for C# and F# teams
Apache Spark's native interfaces are Scala, Java, Python and R. A .NET shop that wants to process data at Spark scale has historically had to run a second stack: a Python or Scala service alongside the C# application, with the serialization and deployment overhead that implies. .NET for Apache Spark exists to remove that split. The README describes it as providing high performance APIs for using Apache Spark from C# and F#, covering the DataFrame and SparkSQL surfaces for structured data and Structured Streaming for streaming data.
The target reader is a .NET developer who already knows the ecosystem and wants to reuse existing skills, code and libraries rather than learn PySpark. That is a real constraint for teams whose production services, build pipelines and dependency management are all .NET. The project is .NET Standard compliant, which the README frames as the reason it can be used anywhere you write .NET code. It runs on Windows, Linux and macOS using .NET 8, or Windows using .NET Framework, and the README lists Azure HDInsight Spark, Amazon EMR Spark and Databricks on both AWS and Azure as supported deployment targets.
One thing to be clear about: this is a binding layer, not a distribution. Spark itself still runs on the JVM. The .NET process talks to it. That distinction drives most of the trade-offs below.
How the .NET to JVM bridge actually moves your data
The architecture implied by the repository and README is a two-process model. Your application is a .NET assembly. Spark is a JVM process. The .NET side calls into the Spark APIs through an interop layer, and the repository layout reflects that split: src/ holds the .NET libraries, while build.sh, build.cmd, eng/, script/ and the Azure Pipelines YAML files handle producing and shipping both halves. The samples under examples/ are split into Microsoft.Spark.CSharp.Examples and Microsoft.Spark.FSharp.Examples, which tells you the API surface is designed to be idiomatic from both languages rather than C#-only with F# as an afterthought.
The supported version matrix is the part to read carefully. The README's table maps .NET for Apache Spark v2.3.1 to Apache Spark 2.4, 3.0, 3.1, 3.2 and 3.5, with a footnote that 2.4.2 is not supported. Note the gap: 3.3 and 3.4 are not listed. If your cluster runs one of those, the README does not claim support, and you should treat that as unverified rather than assume it works.
The API coverage is deliberately partial. The README names DataFrame and SparkSQL for structured data and Structured Streaming for streaming. It does not claim MLlib, GraphX or the low-level RDD API. If your pipeline depends on those, this binding is the wrong layer. There is also an open Spark Project Improvement Proposal, SPARK-27006, referenced in the README, aimed at getting .NET support into Apache Spark by default. Until that lands, .NET support is an external project with its own release cadence, not part of Spark itself.
Installing .NET for Apache Spark and running a first DataFrame job
The README points to platform-specific getting-started guides for Windows, Ubuntu and macOS, all targeting .NET 8. The NuGet package is Microsoft.Spark. The distribution model is a NuGet reference plus a Spark installation on the machine or cluster, so the install is two steps rather than one.
The README gives this example of adding the package, and the package name is the one its badge links to:
dotnet add package Microsoft.SparkAfter the restore completes, your project references Microsoft.Spark and its transitive dependencies. The next step is the entry point. The repository's getting-started samples under examples/Microsoft.Spark.CSharp.Examples/Sql/Batch/ are the reference for the shape of a batch job, and the README links StructuredNetworkWordCount.cs and StructuredNetworkWordCountWindowed.cs under Sql/Streaming for the streaming case, with F# equivalents alongside them. Those are the files to read before writing your own streaming job, because session lifetime and output-mode handling differ from the batch case.
Build status for the project is tracked in Azure Pipelines for Ubuntu and Windows. If you build from source, the README points at build.sh and build.cmd and estimates the full clone-to-running-app process at under 15 minutes, with separate instructions for Windows .NET Framework 4.8, Windows .NET 8 and Ubuntu .NET 8.
Where the version matrix and API surface will bite you
The clearest limitation is the Spark version ceiling. The support table stops at 3.5. There is no entry for Spark 4, and no entry for 3.3 or 3.4. A team on a managed platform that has already moved past 3.5 has no documented path here. That is not a bug report, it is a statement about what the README claims, and it should drive your first check before any adoption decision.
The second limitation is API coverage. DataFrame, SparkSQL and Structured Streaming are in. MLlib, GraphX and RDD-level operations are not mentioned as supported. A machine-learning pipeline that needs MLlib from .NET has no documented route in this repository, even though the topics list includes machine-learning and the samples include TPC-DS and TPC-H benchmark scenarios under benchmark/. Benchmarks measure query performance, not model training.
The third is the bridge itself. Any interop layer between a .NET process and a JVM adds a failure surface that a pure PySpark or Scala job does not have: two runtimes to version-match, two sets of logs, and startup errors that surface at session creation. The README does not document rollback procedures or a compatibility shim for mismatched versions, so if you deploy a .NET for Apache Spark job against a cluster whose Spark version is not in the table, you are outside what the project documents.
PySpark and Scala as the alternatives, and the real difference
The honest alternative is PySpark. It is part of Apache Spark itself, ships with every Spark release, and gets new APIs at the same time as Scala. The difference in approach is not performance, it is ownership: PySpark is maintained inside the Spark project, while .NET for Apache Spark is an external binding whose version support table is set by its own maintainers. If your team can write Python, PySpark removes the version-matrix problem entirely, because the Spark version you run is the version you get. The cost is that you now maintain Python code alongside your .NET services, which is exactly the split this project was built to avoid.
The second alternative is Scala or Java, using Spark's native API. You get complete API coverage including MLlib and RDDs, and no bridge. The cost is a second language in the stack and a JVM build pipeline. For a .NET-heavy organization, neither alternative is free; the question is whether you would rather pay in language diversity or in version lag and API gaps. .NET for Apache Spark is the right answer only when the .NET constraint is firm and the Spark version is inside the table.
Maintenance cadence, licence and what upgrading costs
The repository is not archived, and the last push was on 2026-09-09, which is recent. The release history is slower than Spark's own: v2.3.0 shipped in May 2025, v2.3.1-rc1 in February 2026, and v2.3.1 in February 2026. That cadence matters because Spark releases on its own schedule, and each new Spark minor version is a potential support-table update here. A team that upgrades Spark frequently will spend time waiting on this project rather than on their own code.
The licence is MIT, stated in the repository's LICENSE file. MIT is permissive: it allows use, modification and redistribution with the copyright notice and licence text retained. That is a statement about the licence text, not legal advice; if you redistribute the library inside a product, your legal team should confirm notice requirements. The repository also carries a THIRD-PARTY-NOTICES.TXT, which is where the transitive dependency attributions live, and NuGet.config at the root controls package sources for the build.
Upgrade cost has two dimensions. On the .NET side, the README targets .NET 8, so a team on an older runtime has a migration ahead of it. On the Spark side, every cluster upgrade needs a check against the support table before it happens, because the table is the only documented statement of compatibility in this repository.
Editorial conclusion
Adopt .NET for Apache Spark if your team writes C# or F# and the data platform is already Apache Spark 2.4 through 3.5, particularly on HDInsight, EMR or Databricks where the deployment docs cover the setup. Do not adopt it if you need Spark 4, if you need MLlib, GraphX or RDD-level APIs from .NET, or if your team is comfortable in Scala or PySpark, since those get new Spark features first. Before committing, verify the exact Spark version your cluster runs against the support table, confirm the Microsoft.Spark NuGet package version you intend to pin, and check whether the deployment guide for your specific cloud covers your configuration.
Frequently asked questions
What is .NET for Apache Spark?
It is a set of .NET Standard libraries that expose Apache Spark's DataFrame, SparkSQL and Structured Streaming APIs to C# and F#. It runs alongside a JVM Spark process rather than replacing it.
Which Apache Spark versions does .NET for Apache Spark support?
The README's support table maps v2.3.1 to Apache Spark 2.4, 3.0, 3.1, 3.2 and 3.5, with a note that 2.4.2 is not supported. Spark 3.3, 3.4 and 4 are not listed.
How do I install .NET for Apache Spark?
Add the Microsoft.Spark NuGet package to your project and install Apache Spark separately on the machine or cluster. The README links to getting-started guides for Windows, Ubuntu and macOS, all using .NET 8.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/dotnet-spark)