h2non/filetype: Go type detection from magic numbers, with no dependencies
Fast, dependency-free Go package to infer binary file types based on the magic numbers header signature
At a glance
- What is it?
- A small Go package that infers file type from the header bytes of a file rather than trusting its name, mapping to both an extension and a MIME type. Useful as a library inside a service, less useful as a command line tool.
- Who is it for?
- This package earns its place when you are validating uploads inside a Go service and cannot trust the filename a client sent, since it reads the bytes rather than the name, needs no cgo and no external tool, and lets you add a matcher for a private format in a few lines.
- Can I use it commercially?
- Yes. MIT is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 98 days ago.
- What is it written in?
- Mainly Go, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 28, 2026, and from our analysis. They are not legal advice.
Editorial analysis
Reading the bytes instead of believing the filename
The premise is narrow and correct: a file's name and extension are claims, while the first few dozen bytes are evidence. The package reads a header and matches it against a table of magic number signatures to produce a type object carrying both an extension and a MIME type. Because the check is on content, a `.jpg` that is actually a PNG comes back as PNG, which is exactly the case a server handling uploads needs to get right.
The install is one line, because there is nothing to compile beyond your own program:
go get github.com/h2non/filetypeThe README stresses dependency free in the sense that matters for Go services: just Go code, with no C compilation and therefore no cgo. `go.mod` declares the module at `go 1.13`, so the package drops into almost any Go codebase without raising a toolchain floor, although the README badge advertises Go 1.0 and up, which is a marketing claim the module file quietly contradicts.
The README also points at two sibling projects for the cases this one does not cover: `go-is-svg` for SVG detection, and `filetype.py` as a Python port. Those pointers are worth noting because they reveal where the author considers the boundaries of the approach to be. SVG has no reliable magic number at the start of the file, which is the real reason it is a separate package rather than a missing feature.
The API is four functions wide
The basic call reads a buffer and returns a kind. `filetype.Match` gives you the type, and `filetype.Unknown` is the sentinel when nothing matched, which means the error case has a normal value rather than an exception path:
buf, _ := os.ReadFile("sample.jpg")
kind, _ := filetype.Match(buf)
if kind == filetype.Unknown {
fmt.Println("Unknown file type")
return
}
fmt.Printf("File type: %s. MIME: %s\n", kind.Extension, kind.MIME.Value)Above that sit class helpers, which is the part that saves the most typing in a validation path. `filetype.IsImage`, `filetype.IsVideo` and the audio equivalent answer the question you usually actually have, and the same applies to `filetype.IsSupported` and `filetype.IsMIMESupported` when you need to know whether a type exists in the tables at all rather than what a specific file is.
Two subpackages are part of the public surface: `types` and `matchers`. That split is the package's main extension point, and it is what separates this from a hardcoded switch statement. The repository tree is correspondingly small and flat: `filetype.go`, `match.go`, `kind.go`, `version.go`, a test file beside each, plus `fixtures/`, `matchers/` and `types/`. There is no cmd directory and no binary, which is the clearest statement of what this is: a library, not a tool.
Only the header is read, and there is a rounding error in the docs
The feature list says you need only the first 262 bytes, the maximum file header. The header example in the README then allocates 261 and the comment beside it says first 261 bytes. The two disagree by one byte, which in practice changes nothing, since both are far more than any matcher needs, but it is the kind of detail that tells you the documentation was written by hand and not generated.
The practical point is that you never have to read a whole file. For an upload handler this is the difference between accepting a 4 GB video by reading its header and buffering the entire body:
file, _ := os.Open("movie.mp4")
head := make([]byte, 261)
file.Read(head)
if filetype.IsImage(head) {
fmt.Println("File is an image")
} else {
fmt.Println("Not an image")
}The same reasoning applies to a stream, a network connection or an `http.Request` body. Type detection becomes a constant cost regardless of file size, which is what lets you validate before deciding how to handle the rest of the payload.
One caveat worth stating plainly: this approach only works for formats that begin with a recognizable signature. Every entry in the type tables is a binary container with a magic number at a known offset. Text formats have no such marker, so the absence of a match on an HTML or CSV upload means nothing was recognised, not that the file is invalid. Anything content-based needs different handling.
Adding a matcher for a format nobody else ships
Private and internal formats are the reason to pick this package over anything with a fixed table, and the extension path is short. You declare a type with an extension and MIME string, write a function that inspects the buffer, and register the pair:
var fooType = filetype.NewType("foo", "foo/foo")
func fooMatcher(buf []byte) bool {
return len(buf) > 1 && buf[0] == 0x01 && buf[1] == 0x02
}
filetype.AddMatcher(fooType, fooMatcher)After registration the new type participates in the normal API, so `IsSupported`, `IsMIMESupported` and `Match` all see it. Nothing else has to change, and the built-in matchers live in the `matchers/` subpackage, so the convention is visible by reading the directory.
This is the design that separates filetype from a lookup map you write yourself in twenty lines. Writing your own costs more than twenty lines once you care about the class helpers, the consistent unknown value, and the fact that a matcher list stays in one place. It also means the standard detection logic you did not write is the part already covered by the test files sitting next to each source file.
What is in the tables, and what is missing
The supported list is grouped by class and totals roughly a hundred entries. Image covers jpg, png, gif, webp, cr2, tif, bmp, heif, jxr, psd, ico, dwg and avif, which is a wider set than most libraries manage because cr2 and psd are raw camera and layered image formats rather than web assets. Video covers mp4, m4v, mkv, webm, mov, avi, wmv, mpg, flv and 3gp. Audio covers mid, mp3, m4a, ogg, flac, wav, amr, aac and aiff.
The Archive group is the interesting one, since it is doing double duty as an executables-and-packages group: zip, tar, rar, gz, bz2, 7z, xz, zstd, epub, pdf, iso, cab, deb, rpm, ar, lz, Z, eot, ps, rtf, swf, plus exe, elf, wasm, dex, dey, crx, nes and dcm. Under Documents sit the six office formats from doc through pptx. Font has woff, woff2, ttf and otf, and Application has wasm, dex, dey and parquet.
Two gaps follow from the design rather than from neglect. There is no SVG, for the reason given earlier, and no text or markup formats at all, because those have no magic number to match. A file with no extension and no signature is `Unknown` in both cases, and the API gives you no way to distinguish unknown-binary from known-text-format. Worth knowing before you write a validator that assumes a positive result means the file is well formed.
A release from 2021 and commits that did not stop
The version picture deserves a note because it looks contradictory at first. The only recorded release is v1.1.1, published on 2021-01-21, while the last push to the repository was on 2026-07-01. The repository is not archived and the default branch is `master`.
So the project is being worked on without cutting tagged releases. For a library consumed through `go get` that matters less than it would for a binary, because Go resolves the module to a commit rather than to a published tag, and a consumer tracking `master` picks up the recent commits. It also means there is no changelog tied to a version, and the `History.md` file in the tree is the only release history kept in the repository.
On benchmarks, the README publishes figures measured on real files from the `fixtures/` directory on an OSX x64 i7 2.7 GHz machine, with the tar, zip and jpeg match benchmarks each running a million iterations in roughly 1.1 to 1.3 microseconds. That is old hardware by current standards and the numbers should be read as relative rather than absolute, but the shape is the point: header matching is a sequence of short byte comparisons over about a hundred matchers, and at that size the cost is dominated by the loop rather than by any I/O. Note that the published benchmark block is cut off mid-line, so the jpeg figure has no complete unit in the README.
Under MIT there is nothing unusual to weigh for internal use. The licence obligation that matters to a company is the copyright notice retention in redistributed binaries, which for a linked Go library means your own distribution, not this package's.
Editorial conclusion
This package earns its place when you are validating uploads inside a Go service and cannot trust the filename a client sent, since it reads the bytes rather than the name, needs no cgo and no external tool, and lets you add a matcher for a private format in a few lines. It is the wrong choice for a desktop user who just wants to identify one file, because it ships as a library rather than a command, and it is the wrong choice for text formats, where the header carries no signature at all. Check two things before adopting it: whether your formats have a real magic number, since the type tables cover binary containers mainly, and whether the gap since v1.1.1 matters to you, given the last commit is far newer than the last tag.
Frequently asked questions
How do I install the filetype Go package?
With `go get github.com/h2non/filetype`. There is nothing else to install, because the package is dependency free Go code with no C compilation and no cgo, and `go.mod` declares Go 1.13. It is a library rather than a command line binary, so there is nothing to add to your PATH.
How many bytes does filetype need to identify a file?
Only the header. The feature list states 262 bytes as the maximum file header, and the header example in the README allocates a 261 byte slice and reads just that much. The one byte discrepancy between the two does not matter in practice, since both are far beyond what any matcher inspects.
Does filetype detect SVG, HTML or other text formats?
SVG is not supported, and the README points at a separate `go-is-svg` package for it, because SVG has no magic number signature at the start of the file. Text and markup formats are likewise absent for the same reason. A match returning Unknown therefore means no signature was recognised, not that the file is malformed.
Can I add support for my own private file format?
Yes. Declare a type with `filetype.NewType`, write a matcher function that inspects the buffer, and register the pair with `filetype.AddMatcher`. The new type then participates in `Match`, `IsSupported` and `IsMIMESupported` like a built-in one, and built-in matchers live in the `matchers/` subpackage to serve as a reference.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/h2non-filetype)