darold/pgbadger: PostgreSQL logs turned into a report you can act on
A fast PostgreSQL Log Analyzer
At a glance
- What is it?
- A single Perl program that reads Postgres server logs and produces detailed HTML reports with graphs, built around twenty years of accumulated parsing edge cases.
- Who is it for?
- pgBadger does one job well and has kept doing it long enough that the parsing rules are the interesting part. The option list reads like a history of everything that has ever gone wrong in a Postgres log file, from multiline statements that produce enormous error dumps to placeholder parsing mistakes in log_line_prefix.
- Can I use it commercially?
- Yes. PostgreSQL is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 21 days ago.
- What is it written in?
- Mainly Perl, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 23, 2026, and from our analysis. They are not legal advice.
Editorial analysis
One Perl program that reads logs and writes an HTML report
pgBadger is a Perl script with a Perl distribution wrapper around it. The invocation is short enough to fit in your head:
Usage: pgbadger \[options\] logfile \[...\]The argument is deliberately flexible. A single log file works, a list of files works, and a shell command that returns a list of files works too, which is how you would point it at a directory of rotated logs without building the list yourself. Passing `-` reads from standard input, with the documented caveat that stdin does not work with csvlog.
The repository layout reflects a long-lived CPAN style project rather than an application. `pgbadger` is the script itself, `Makefile.PL` and `META.yml` are the distribution build files, `MANIFEST` lists what ships, and there are three directories worth knowing: `doc/` for the extended documentation, `t/` for the test suite, and `tools/` for helper scripts. There is also `HACKING.md` and `CONTRIBUTING.md`, and a `ChangeLog` that has been kept for a long time.
The license is the PostgreSQL licence, which is permissive and business friendly. At 4065 stars and 377 forks, this is one of the better known third party tools in the Postgres ecosystem, and the repository is not archived, with the last push recorded on 2026-09-15 and 23 open issues.
The option list is where the real documentation lives
The README leads with a full synopsis rather than a tutorial, and that turns out to be the right choice for a tool with this many input shapes. The options that matter first are the parsing ones. `-f` forces a log format when auto-detection is not confident, and the accepted values are syslog, syslog2, stderr, jsonlog, csv, pgbouncer, logplex, rds and redshift. Those last three matter more than they look: pgBouncer logs, Rackspace-style logplex output and Amazon Redshift or RDS log formats are all non-standard sources that pgBadger will happily parse if you tell it what it is looking at.
The prefix option is the subtle one. `-p` takes a custom log_line_prefix value for when you are not using one of the standard prefixes, and the documentation states the requirement plainly: it must contain escape sequences for time (`%t`, `%m` or `%n`) and processes (`%p` or `%c`). That constraint is not bureaucratic. Without a timestamp and a process identifier there is no way to group or order events, so the parser needs both.
The output side is equally explicit. `-x` selects text, html, bin or json, and `-o` names the file, defaulting to out.html, out.txt, out.bin or out.json to match. JSON output requires the `JSON::XS` Perl module, and `-o -` dumps to stdout. Filters for narrowing the report are all there too: `-d` for database, `-u` for user, `-U` to exclude a user, `-c` for client host, `-N` for application name, `-S` to report only SELECT queries, and `-D` to replace client addresses with DNS names.
Two kinds of parallelism, and a switch to turn both off
pgBadger separates parallelism within a file from parallelism across files, which matters when you are deciding how to run it against a large archive. `-j` sets the number of jobs to run at the same time on a single log file, and it defaults to a single job, and to a single job automatically when working with csvlog. That last part is a real constraint rather than a preference, because the CSV format is positional and does not split as freely.
`-J` is the other axis: the number of log files to parse in parallel, where the default is to process one file at a time. Running both together is the configuration for a multi-gigabyte archive on a machine with cores to spare:
pgbadger -j 8 -J 4 -t 10`-t` sets how many queries are stored and displayed, defaulting to 20, and `-s` sets how many query samples are kept per entry, defaulting to 3. Those two numbers control report size more than anything else, since a busy server produces thousands of distinct query texts.
For debugging, version 13.1 added `--no-fork`, which stops the process forking at all. That is the flag you want when you need to attach a debugger or watch the parse happen line by line, and its presence tells you the maintainer expects people to need that.
Incremental mode for logs that never stop arriving
Long lived servers keep producing logs, and regenerating a report over months of history on every run is wasteful. The incremental mode exists for that. `-I` switches it on and requires `--outdir` to be set, after which reports are generated by day into a separate directory:
pgbadger -I -O out --html-outdir out`-H` sets the directory for HTML reports in incremental mode specifically, because the binary intermediate files stay in the directory given by `-O`. `-R` sets a retention count in weeks, defaulting to zero for disabled, and older week and day directories are removed automatically. `-l` records the last datetime and line parsed, which is what makes it possible to resume rather than reparse, and it doubles as an error watcher when you want to see only what happened since the last run.
`-L` takes a file containing a list of log files to parse, `-w` restricts output to errors only, the way logwatch would, and `-E` explodes the main report into one report per database, folding global information into the postgres database report. `-a` controls the averaging window in minutes, defaulting to 5, and `-A` controls the histogram window, defaulting to 60. Getting these two wrong is the usual reason a report's graphs look empty over a long time range.
What the release notes reveal about parser edge cases
The release history is more informative about what Postgres logs actually look like than the README is. Three releases are tagged, and each is a maintenance release built from user reports rather than a rewrite.
Version 13.2, published 2025-12-29, fixes normalization when single-quoted strings contain escaped quotes, fixes a precedence problem between `!` and `%s`, fixes parsing of the `%r` placeholder in log_line_prefix, updates the bundled pgFormatter to version 5.9, and adds `--ssh-sudo` for running commands over ssh as sudo. It also adds a GitHub CI action for testing on commit push, which is the sort of thing that should have existed years earlier.
Version 13.1, from 2025-03-17, adds a vacuum throughput report with a graph of vacuum per table, including per-table I/O timing for reads and writes plus elapsed CPU time, and adds frozen pages and tuple counts to it. It also adds milliseconds to the raw CSV output, records the log filename in sample reports when several files are processed, and fixes bind parameter parsing and query filtering on multiline queries.
Version 13.0, from 2024-12-08, is the one release that changes output structure. It adds `--histogram-query` and `--histogram-session` for custom histogram boundaries, supports auto_explain plans in CSV and JSON log formats, and reports three LOG level messages that were previously missed: unexpected EOF, incomplete startup packet and deadlock detected while waiting. That release also notes a backward compatibility break, since the way LOG level events are stored in the Events reports changed. Scripts reading the binary or JSON output need to know that.
Where it sits against the alternatives
The honest comparison is not with other log analyzers but with running queries against the database itself. Postgres already collects view statistics, and `pg_stat_statements` will tell you which normalised queries are slow. What pgBadger does that those cannot is show you errors, connection events, checkpoints, autovacuum activity and temporary file usage on one timeline, from a log file you can replay at any time after the fact.
That replayability is the strongest argument for keeping logs at all. A statistics view reflects the current state of shared memory and tells you nothing about the incident that happened yesterday afternoon. A log file, parsed after the fact, does.
The counterweight is parsing fragility. Log format is a configuration setting on the server, and if someone changes log_line_prefix or switches between stderr and csvlog, the report changes shape. The `-p` and `-f` options exist because of exactly this, and using csvlog as the destination removes most of the ambiguity because every field is labelled. The project also pulls in external pieces worth knowing about: the bundled pgFormatter handles query prettification and can be disabled with `-P`, and comment stripping with `-C` or disabling multiline collection with `-M` are both there for when noisy queries would otherwise dominate the report. Version 13.2 shipping a pgFormatter update to 5.9 is a reminder that this dependency moves on its own schedule.
Editorial conclusion
pgBadger does one job well and has kept doing it long enough that the parsing rules are the interesting part. The option list reads like a history of everything that has ever gone wrong in a Postgres log file, from multiline statements that produce enormous error dumps to placeholder parsing mistakes in log_line_prefix. Two practical takeaways: give it csvlog rather than stderr when you can, because the CSV format removes most of the guessing, and use the incremental mode with an explicit outdir if your logs rotate, so a month of history stays browsable. Start with the SYNOPSIS section in the README, generate one HTML file, then read the docs directory and the ChangeLog to see which report suits your question. The homepage at pgbadger.darold.net carries the fuller manual for configuration and sample output.
Frequently asked questions
How do you install pgBadger?
It is a Perl program, so the usual route is a CPAN install or building from the repository with Makefile.PL followed by make and make install. The README here starts at the synopsis rather than the install section, and the extended documentation in the doc directory plus the homepage at pgbadger.darold.net carry the full instructions.
What log formats can pgBadger parse?
The `-f` option lists syslog, syslog2, stderr, jsonlog, csv, pgbouncer, logplex, rds and redshift. The pgbouncer, logplex and rds entries cover non-Postgres sources such as connection pooler logs and managed database log formats. Force the format with `-f` whenever detection is not confident.
Can pgBadger produce a report for a custom log_line_prefix?
Yes, through the `-p` option, with one requirement: the prefix string must contain a time escape sequence, either `%t`, `%m` or `%n`, and a process escape sequence, either `%p` or `%c`. Without both, pgBadger has no way to order or group the entries it reads.
How do you make pgBadger faster on large log files?
Two separate options apply. `-j` sets parallel jobs within a single log file and stays at one job when the format is csvlog, while `-J` parses that many log files at once. Running both together is the usual answer, and `--no-fork` added in version 13.1 helps when you need to watch what the parser is doing.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/darold-pgbadger)