# h2non/filetype: Go type detection from magic numbers, with no dependencies

> A small Go package that infers file type from the header bytes of a file rather than trusting its name, mapping to both an extension and a MIME type. Useful as a library inside a service, less useful as a command line tool.

**h2non/filetype** — Fast, dependency-free Go package to infer binary file types based on the magic numbers header signature

- Repository: https://github.com/h2non/filetype
- Website: https://pkg.go.dev/github.com/h2non/filetype?tab=doc
- Stars: 2,301 · Forks: 190
- Language: Go
- License: MIT
- Published: 2026-10-06 · Updated: 2026-10-06 · Language: en
- Canonical page: https://hysenlabs.com/projects/h2non-filetype

## Reading the bytes instead of believing the filename

The premise is narrow and correct: a file's name and extension are claims, while the first few dozen bytes are evidence. The package reads a header and matches it against a table of magic number signatures to produce a type object carrying both an extension and a MIME type. Because the check is on content, a `.jpg` that is actually a PNG comes back as PNG, which is exactly the case a server handling uploads needs to get right.

The install is one line, because there is nothing to compile beyond your own program:

```bash
go get github.com/h2non/filetype
```

The README stresses dependency free in the sense that matters for Go services: just Go code, with no C compilation and therefore no cgo. `go.mod` declares the module at `go 1.13`, so the package drops into almost any Go codebase without raising a toolchain floor, although the README badge advertises Go 1.0 and up, which is a marketing claim the module file quietly contradicts.

The README also points at two sibling projects for the cases this one does not cover: `go-is-svg` for SVG detection, and `filetype.py` as a Python port. Those pointers are worth noting because they reveal where the author considers the boundaries of the approach to be. SVG has no reliable magic number at the start of the file, which is the real reason it is a separate package rather than a missing feature.

## The API is four functions wide

The basic call reads a buffer and returns a kind. `filetype.Match` gives you the type, and `filetype.Unknown` is the sentinel when nothing matched, which means the error case has a normal value rather than an exception path:

```go
buf, _ := os.ReadFile("sample.jpg")
kind, _ := filetype.Match(buf)
if kind == filetype.Unknown {
  fmt.Println("Unknown file type")
  return
}
fmt.Printf("File type: %s. MIME: %s\n", kind.Extension, kind.MIME.Value)
```

Above that sit class helpers, which is the part that saves the most typing in a validation path. `filetype.IsImage`, `filetype.IsVideo` and the audio equivalent answer the question you usually actually have, and the same applies to `filetype.IsSupported` and `filetype.IsMIMESupported` when you need to know whether a type exists in the tables at all rather than what a specific file is.

Two subpackages are part of the public surface: `types` and `matchers`. That split is the package's main extension point, and it is what separates this from a hardcoded switch statement. The repository tree is correspondingly small and flat: `filetype.go`, `match.go`, `kind.go`, `version.go`, a test file beside each, plus `fixtures/`, `matchers/` and `types/`. There is no cmd directory and no binary, which is the clearest statement of what this is: a library, not a tool.

## Only the header is read, and there is a rounding error in the docs

The feature list says you need only the first 262 bytes, the maximum file header. The header example in the README then allocates 261 and the comment beside it says first 261 bytes. The two disagree by one byte, which in practice changes nothing, since both are far more than any matcher needs, but it is the kind of detail that tells you the documentation was written by hand and not generated.

The practical point is that you never have to read a whole file. For an upload handler this is the difference between accepting a 4 GB video by reading its header and buffering the entire body:

```go
file, _ := os.Open("movie.mp4")
head := make([]byte, 261)
file.Read(head)

if filetype.IsImage(head) {
  fmt.Println("File is an image")
} else {
  fmt.Println("Not an image")
}
```

The same reasoning applies to a stream, a network connection or an `http.Request` body. Type detection becomes a constant cost regardless of file size, which is what lets you validate before deciding how to handle the rest of the payload.

One caveat worth stating plainly: this approach only works for formats that begin with a recognizable signature. Every entry in the type tables is a binary container with a magic number at a known offset. Text formats have no such marker, so the absence of a match on an HTML or CSV upload means nothing was recognised, not that the file is invalid. Anything content-based needs different handling.

## Adding a matcher for a format nobody else ships

Private and internal formats are the reason to pick this package over anything with a fixed table, and the extension path is short. You declare a type with an extension and MIME string, write a function that inspects the buffer, and register the pair:

```go
var fooType = filetype.NewType("foo", "foo/foo")

func fooMatcher(buf []byte) bool {
  return len(buf) > 1 && buf[0] == 0x01 && buf[1] == 0x02
}

filetype.AddMatcher(fooType, fooMatcher)
```

After registration the new type participates in the normal API, so `IsSupported`, `IsMIMESupported` and `Match` all see it. Nothing else has to change, and the built-in matchers live in the `matchers/` subpackage, so the convention is visible by reading the directory.

This is the design that separates filetype from a lookup map you write yourself in twenty lines. Writing your own costs more than twenty lines once you care about the class helpers, the consistent unknown value, and the fact that a matcher list stays in one place. It also means the standard detection logic you did not write is the part already covered by the test files sitting next to each source file.

## What is in the tables, and what is missing

The supported list is grouped by class and totals roughly a hundred entries. Image covers jpg, png, gif, webp, cr2, tif, bmp, heif, jxr, psd, ico, dwg and avif, which is a wider set than most libraries manage because cr2 and psd are raw camera and layered image formats rather than web assets. Video covers mp4, m4v, mkv, webm, mov, avi, wmv, mpg, flv and 3gp. Audio covers mid, mp3, m4a, ogg, flac, wav, amr, aac and aiff.

The Archive group is the interesting one, since it is doing double duty as an executables-and-packages group: zip, tar, rar, gz, bz2, 7z, xz, zstd, epub, pdf, iso, cab, deb, rpm, ar, lz, Z, eot, ps, rtf, swf, plus exe, elf, wasm, dex, dey, crx, nes and dcm. Under Documents sit the six office formats from doc through pptx. Font has woff, woff2, ttf and otf, and Application has wasm, dex, dey and parquet.

Two gaps follow from the design rather than from neglect. There is no SVG, for the reason given earlier, and no text or markup formats at all, because those have no magic number to match. A file with no extension and no signature is `Unknown` in both cases, and the API gives you no way to distinguish unknown-binary from known-text-format. Worth knowing before you write a validator that assumes a positive result means the file is well formed.

## A release from 2021 and commits that did not stop

The version picture deserves a note because it looks contradictory at first. The only recorded release is v1.1.1, published on 2021-01-21, while the last push to the repository was on 2026-07-01. The repository is not archived and the default branch is `master`.

So the project is being worked on without cutting tagged releases. For a library consumed through `go get` that matters less than it would for a binary, because Go resolves the module to a commit rather than to a published tag, and a consumer tracking `master` picks up the recent commits. It also means there is no changelog tied to a version, and the `History.md` file in the tree is the only release history kept in the repository.

On benchmarks, the README publishes figures measured on real files from the `fixtures/` directory on an OSX x64 i7 2.7 GHz machine, with the tar, zip and jpeg match benchmarks each running a million iterations in roughly 1.1 to 1.3 microseconds. That is old hardware by current standards and the numbers should be read as relative rather than absolute, but the shape is the point: header matching is a sequence of short byte comparisons over about a hundred matchers, and at that size the cost is dominated by the loop rather than by any I/O. Note that the published benchmark block is cut off mid-line, so the jpeg figure has no complete unit in the README.

Under MIT there is nothing unusual to weigh for internal use. The licence obligation that matters to a company is the copyright notice retention in redistributed binaries, which for a linked Go library means your own distribution, not this package's.

## Conclusion

This package earns its place when you are validating uploads inside a Go service and cannot trust the filename a client sent, since it reads the bytes rather than the name, needs no cgo and no external tool, and lets you add a matcher for a private format in a few lines. It is the wrong choice for a desktop user who just wants to identify one file, because it ships as a library rather than a command, and it is the wrong choice for text formats, where the header carries no signature at all. Check two things before adopting it: whether your formats have a real magic number, since the type tables cover binary containers mainly, and whether the gap since v1.1.1 matters to you, given the last commit is far newer than the last tag.

## FAQ

### How do I install the filetype Go package?

With `go get github.com/h2non/filetype`. There is nothing else to install, because the package is dependency free Go code with no C compilation and no cgo, and `go.mod` declares Go 1.13. It is a library rather than a command line binary, so there is nothing to add to your PATH.

### How many bytes does filetype need to identify a file?

Only the header. The feature list states 262 bytes as the maximum file header, and the header example in the README allocates a 261 byte slice and reads just that much. The one byte discrepancy between the two does not matter in practice, since both are far beyond what any matcher inspects.

### Does filetype detect SVG, HTML or other text formats?

SVG is not supported, and the README points at a separate `go-is-svg` package for it, because SVG has no magic number signature at the start of the file. Text and markup formats are likewise absent for the same reason. A match returning Unknown therefore means no signature was recognised, not that the file is malformed.

### Can I add support for my own private file format?

Yes. Declare a type with `filetype.NewType`, write a matcher function that inspects the buffer, and register the pair with `filetype.AddMatcher`. The new type then participates in `Match`, `IsSupported` and `IsMIMESupported` like a built-in one, and built-in matchers live in the `matchers/` subpackage to serve as a reference.

## Sources

- [h2non/filetype on GitHub](https://github.com/h2non/filetype)
- [License: MIT](https://github.com/h2non/filetype/blob/master/LICENSE)
- [Project website](https://pkg.go.dev/github.com/h2non/filetype?tab=doc)
- [README](https://github.com/h2non/filetype/blob/master/README.md)
- [Releases](https://github.com/h2non/filetype/releases)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/h2non-filetype
