TrailScraper: turning AWS CloudTrail logs into IAM policies
A command-line tool to get valuable information out of AWS CloudTrail.
At a glance
- What is it?
- TrailScraper is a Python command line tool that reads CloudTrail events and emits IAM policy JSON. It is useful for tightening permissions after the fact, but its action mapping is heuristic and the project labels itself alpha.
- Who is it for?
- Adopt TrailScraper if you already keep CloudTrail logs and want a starting point for least-privilege IAM policies rather than a hand-written guess. Do not adopt it if you need a guaranteed, machine-verified mapping from CloudTrail events to IAM actions, or if your trails do not cover us-east-1.
- Can I use it commercially?
- Yes. Apache-2.0 is a permissive licence: you can use, modify and sell software built on it, as long as you keep its copyright and licence notices.
- Is it still maintained?
- Yes. The repository last received commits 8 days ago.
- What is it written in?
- Mainly Python, according to GitHub's language statistics.
Answers come from the project's GitHub data, last synced on September 25, 2026, and from our analysis. They are not legal advice.
Editorial analysis
What TrailScraper does with CloudTrail data
CloudTrail records what happened in an AWS account. It does not tell you what a role is allowed to do, and it does not produce a policy document. TrailScraper sits in that gap. The README describes it as "a command-line tool to get valuable information out of AWS CloudTrail and a general purpose toolbox for working with IAM policies".
The intended reader is someone who has to write or trim an IAM policy and would rather derive it from observed API calls than from memory. That covers platform engineers tightening a service role, and security engineers reviewing what a CI role actually touched. The tool is not a monitoring product and not a CloudTrail replacement. It reads events you already have, either through the CloudTrail API or from log files you downloaded, and converts them into IAM policy JSON.
The project classifies itself as "Development Status :: 3 - Alpha" in pyproject.toml, which is worth taking at face value. The last push to the repository was on 2025-12-31, and release 0.10.0 carries the same timestamp.
The select, generate and guess pipeline
The design is a Unix-style pipeline of small subcommands. select produces CloudTrail records as JSON. generate reads records on stdin and writes an IAM policy. guess reads a policy and adds statements it thinks you will also need. The README's own example chains the first two: `trailscraper select | trailscraper generate`.
select has two sources. With `--use-cloudtrail-api` it queries the CloudTrail API directly, filtered by `--filter-assumed-role-arn` and a `--from`/`--to` window. Without that flag it reads locally downloaded logs. The time arguments accept natural language, which is dateparser doing the work: the README shows `--from 'one hour ago'` and `--to 'now'`, and elsewhere `--from 'two days ago'`.
download is the piece that fetches logs from S3, taking `--bucket`, `--account-id` and one or more `--region` flags. Organisational trails add `--org-id`. generate then maps each event's eventSource and eventName onto IAM actions and emits statements with `Effect: Allow` and `Resource: ["*"]`.
The weak point is that mapping. The README is direct about it: "there is no good, machine-readable documentation on how CloudTrail events map to IAM actions so TrailScraper is using heuristics". The repository carries an unknown_actions.txt file, which suggests the heuristics have a known list of events they cannot resolve. Treat generated actions as candidates to check, not as output you paste into production.
Installing TrailScraper and generating a first policy
The README gives three installation routes. On macOS, Homebrew. Otherwise pip, with Python >= 3.5 stated in the README, though pyproject.toml requires Python >= 3.10 for the current package. There is also a container image.
pip install trailscraperAfter that, the trailscraper command is on your PATH. The Docker route passes AWS credentials and config through from the host, which is the quickest way to try it without touching your Python environment:
docker run --rm --env-file <(env | grep AWS_) -v $HOME/.aws:/root/.aws ghcr.io/flosell/trailscraper:latestVersions from 0.7.0 onward are on GitHub Container Registry; older ones are on DockerHub, per the README.
For a first real run, pull the logs for one role over a short window. Note the README's instruction to include us-east-1, because global services such as IAM, STS, Route53 and CloudFront log there:
trailscraper download --bucket some-bucket \
--account-id some-account-id \
--region us-east-1 \
--from 'two days ago' \
--to 'now'Then select the events for the role you care about and pipe them straight into generate. The output is IAM policy JSON with a Version of 2012-10-17 and an array of Statement objects:
trailscraper select --filter-assumed-role-arn some-arn \
--from 'one hour ago' \
--to 'now' | trailscraper generateIf the policy looks too narrow, guess adds related actions. The README's example takes a policy containing s3:PutObject and returns one that also allows s3:DeleteObject, s3:GetObject and s3:ListObjects:
cat minimal-policy.json | trailscraper guessThe guess subcommand also accepts `--only`, shown in the README as `cat minimal-policy.json | ./go trailscraper guess --only`, which restricts output to the guessed statements rather than merging them with the input.
Where TrailScraper gets events wrong
Two failure modes are documented, and both matter more than the feature list.
The first is missing events. The README's answer to "Why is TrailScraper missing some events?" is to check that you have logs for us-east-1, because global services write there. If your download call omits that region, a policy generated from the result will silently lack the IAM, STS or Route53 permissions the role used. Nothing in the output flags the gap. You get a plausible policy with holes in it.
The second is invented actions. The README states plainly that some generated actions "are not real IAM actions", a consequence of the heuristic mapping. A policy containing a nonexistent action will not behave the way the JSON suggests, and it will not necessarily fail loudly at apply time. The project asks users who hit a special case to open an issue or submit a pull request, which tells you the coverage is expected to be incomplete.
There is a scope limit too. Resource is always `["*"]` in the examples. TrailScraper reconstructs actions, not resource ARNs, so a policy it produces is action-scoped but not resource-scoped. If your goal is a policy that names specific buckets or tables, this tool gets you part of the way and no further. And the package metadata says Alpha, so the CLI surface should be expected to move between minor versions.
TrailScraper versus writing the policy by hand
The obvious alternative is reading the CloudTrail console or your own log queries and writing the policy yourself, which is what most teams do today. The difference is direction of inference. Manual authoring starts from intent (this job needs to write to one bucket) and you look up the actions. TrailScraper starts from observed calls and works backward, which catches permissions you forgot a service needed and misses permissions a role needs but has not exercised yet.
That asymmetry is the whole argument. A generated policy is a record of what happened during the window you selected, not a specification of what the role requires. Run it over a quiet week and you will under-provision; run it over a busy one and you may capture one-off actions that should not be permanent.
For teams already running policy-generation tooling from infrastructure code, TrailScraper is complementary rather than competing: it can validate that the actions your Terraform or CloudFormation declares match what the role actually called. The README points at cfn-flip for CloudFormation YAML output and iam-policy-json-to-terraform for HCL, and states that TrailScraper itself does not provide either format. So the JSON is the interchange point, and conversion is somebody else's job.
Licence, dependencies and upgrade cost
TrailScraper is Apache-2.0, stated in both the LICENSE file and the pyproject.toml license field. That is a permissive licence with an explicit patent grant, and it imposes no copyleft obligation on the policies you generate. The generated JSON is your data, not a derivative of the tool. None of this is legal advice; if you redistribute the tool inside a product, read the licence text yourself.
The dependency list is short and pinned exactly: boto3, click, toolz, dateparser, pytz and ruamel.yaml. Exact pins mean upgrades are deliberate rather than automatic, and a stale boto3 pin is the most likely thing to age badly, since it tracks the AWS SDK surface. The project also ships a uv.lock and a .tool-versions file, so the maintainer's own environment is reproducible.
The upgrade cost is mostly the Alpha status. Between 0.8.1 in January 2023 and 0.9.1 in March 2025 the version moved one minor step, then to 0.10.0 at the end of 2025. That is a slow cadence, and the CLI's subcommand and flag names are the contract you depend on in scripts. Pin the version in CI and read CHANGELOG.md before moving.
Editorial conclusion
Adopt TrailScraper if you already keep CloudTrail logs and want a starting point for least-privilege IAM policies rather than a hand-written guess. Do not adopt it if you need a guaranteed, machine-verified mapping from CloudTrail events to IAM actions, or if your trails do not cover us-east-1. Before rolling it out, run trailscraper select against one role and one narrow time window, then diff the generated policy against what that role actually needs.
Frequently asked questions
Is AWS CloudTrail free?
The README does not discuss CloudTrail pricing. It only describes how TrailScraper reads CloudTrail events and downloads logs from S3 buckets, so cost questions have to be answered from AWS's own documentation.
What is a CloudTrail?
The README treats CloudTrail as the source of the events TrailScraper consumes: it records AWS API activity, and TrailScraper selects those records by role ARN and time window or downloads them from a bucket. The README does not define the service itself.
Why is TrailScraper missing some CloudTrail events?
The README points at region coverage: some global AWS services such as IAM, STS, Route53 and CloudFront log to us-east-1, so a download that omits that region will not contain their events.
Can TrailScraper output CloudFormation YAML or Terraform HCL?
No. The README states that TrailScraper does not provide either format, and suggests piping the generated JSON through cfn-flip for CloudFormation or iam-policy-json-to-terraform for HCL.
Why does TrailScraper generate actions that are not real IAM actions?
Because there is no good machine-readable mapping from CloudTrail events to IAM actions, so the tool uses heuristics that the README says likely do not cover all special cases. Users who find an uncovered case are asked to open an issue or submit a pull request.
Official sources
Add this badge to your README
If you maintain this project, the badge below links readers to this analysis and shows its maintenance status from the daily GitHub snapshot. Paste the markdown into your README; add ?metric=license or ?metric=stars to the image URL for a different field.
[](https://hysenlabs.com/projects/flosell-trailscraper)