# activerecord-import: bulk inserts that follow your associations

> activerecord-import adds a single import method to ActiveRecord models so a parent-child-grandchild graph lands in three SQL statements instead of millions. It is a Ruby gem for Rails teams whose batch jobs are dominated by insert round trips, and the README is explicit about which adapters and which duplicate-key clauses are supported.

**zdennis/activerecord-import** — A library for bulk insertion of data into your database using ActiveRecord.

- Repository: https://github.com/zdennis/activerecord-import
- Website: http://www.continuousthinking.com
- Stars: 4,154 · Forks: 620
- Language: Ruby
- License: MIT
- Published: 2026-09-23 · Updated: 2026-09-23 · Language: en
- Canonical page: https://hysenlabs.com/projects/zdennis-activerecord-import

## The N+1 insert problem activerecord-import was built to remove

The gem's own worked example is the clearest statement of its purpose. Given publishers that have books, and books that have reviews, inserting 100 publishers with 10,000 books each and 3 reviews per book through ordinary ActiveRecord saves produces roughly 4M SQL insert statements: 100 for the publishers, then 100 * 10,000 for the books, then 100 * 10,000 * 3 for the reviews. The README states that activerecord-import follows the associations down and generates 3 statements instead, one per table. It also reports that in one case this converted an 18 hour batch process into under 2 hours.

That is the audience: Ruby and Rails teams running data loads, imports from upstream systems, or backfills where the row count is large and the object graph is shallow but wide. The gem is not a general ETL framework and it does not replace your schema. It is a narrower tool that answers one question: how do I get this array of records into the database without paying a round trip per record.

## How import turns an array of models into a few INSERT statements

The gem adds an import method to ActiveRecord classes, with bulk_import as an alias for compatibility with gems such as elasticsearch-model. import accepts several input shapes, and the README ranks them by speed: raw columns plus arrays of values is fastest, model objects are next, and validation adds cost on top of either.

When you pass models, the attributes are pulled off each model by looking at the columns available on the model. When you pass an array of column names alongside the records, those column names decide which fields are written, so Book.import [:title], books writes only the title column and leaves author NULL. Batching is controlled by a batch_size option that sets how many rows go into each INSERT statement; the default is the total number of records, meaning a single statement.

The recursive path is the distinguishing mechanism. Instead of treating each table separately, the gem walks the associations it is given and emits the minimal set of insert statements, which is where the 3-statements-for-three-tables claim comes from. Validations run before the insert when validate is true, and it defaults to true, so the fast path is opt-in: you pass validate: false deliberately.

## Installing activerecord-import and running a first import

The README's requiring section covers autoloading via Bundler and manual loading, so the normal path is adding the gem to your Gemfile and letting Bundler require it. The repository's Dockerfile shows the project's own development setup, not an application install: it builds from ruby:3.2-bullseye, sets AR_VERSION to 7.0, copies test/database.yml.sample to test/database.yml and runs bundle install.

Once the gem is loaded, the smallest real import uses columns and arrays, which the README calls the fastest and most primitive mechanism. The columns array names the target columns in order and each inner array is one record's values in that same order.

```ruby
columns = [ :title, :author ]
values = [ ['Book1', 'George Orwell'], ['Book2', 'Bob Jones'] ]

Book.import columns, values, validate: false
```

With validate: false the gem skips model validations; the README notes that when validate is not specified it defaults to true. If you prefer to work with model objects, build them and pass the array, which is the shape most Rails code already has.

```ruby
books = []
10.times do |i|
  books << Book.new(name: "book #{i}")
end
Book.import books    # or use import!
```

The README states this produces 1 SQL call where the equivalent 10.times { Book.create! ... } loop produces 10. For a larger load, add batch_size to split the work into several INSERT statements rather than one very large one. The repository also ships a docker-compose.yml with mysql:8.0 on port 3306 and postgis/postgis:10-2.5 on port 5432 for running the gem's own test suite.

## Hashes with inconsistent keys raise, and other limits worth knowing

The most concrete failure mode in the README concerns arrays of hashes. Using hashes only works if the columns are consistent in every hash of the array; if they are not, an exception is raised. The README points to issue 507 for discussion and gives two workarounds: instantiate ActiveRecord objects from the hashes and import those, or group the array by its keys and import each group separately.

```ruby
arr = [
  { bar: 'abc' },
  { baz: 'xyz' },
  { bar: '123', baz: '456' }
]

# An exception will be raised
Foo.import arr

# better
arr.map! { |args| Foo.new(args) }
Foo.import arr
```

There is a second boundary around duplicate-key handling. On duplicate key updates require MySQL, SQLite 3.24.0+ or Postgres 9.5+, so the feature is not adapter-neutral. The README also has a callbacks section, and it states that callbacks are not run, which means any before_save or after_create logic in your models will not execute during an import. If your application depends on those hooks for auditing, cache invalidation or derived columns, import is the wrong entry point and you either move that logic into the caller or stay with ordinary saves. Finally, validation is on by default, so the fastest documented path requires you to turn it off consciously and accept that invalid rows reach the database.

## activerecord-import versus insert_all and upsert_all

Rails ships its own bulk insert through insert_all and upsert_all, and that is the natural comparison. The difference is scope. insert_all and upsert_all operate on a single table: you hand them an array of attribute hashes and they build one statement for that table. activerecord-import's headline feature is association traversal, so one call can insert publishers, their books and those books' reviews in three statements. If your load is a single flat table, insert_all is already in the framework and needs no dependency.

On duplicate-key behaviour the two overlap more than people assume. activerecord-import offers duplicate key ignore and duplicate key update options, and the README documents on_duplicate_key_update requiring MySQL, SQLite 3.24.0+ or Postgres 9.5+. upsert_all covers the same ground on modern Rails. The practical split is graph shape, not upsert syntax: reach for activerecord-import when the records form a parent-child hierarchy you would otherwise insert table by table, and stay with the framework when they do not.

## Maintenance, licence and what an upgrade costs you

The repository is not archived, and the last push was on 2026-08-28, so there is recent activity on master. The README top-level entries include a CHANGELOG.md, which is where you should look before bumping the gem, and the repository carries .rubocop.yml and .rubocop_todo.yml, meaning style enforcement is part of the contribution flow.

The licence is MIT, which permits commercial use and modification; the LICENSE file is at the repository root. That is a permissive arrangement, not legal advice, and if your organisation has a policy on third-party licences you should have it applied to the actual LICENSE file rather than to this summary.

The upgrade cost is dominated by adapter and Rails version drift rather than by the gem's own API. The repository's Dockerfile pins AR_VERSION to 7.0 and the docker-compose app service sets the same value, while the gemfiles/ directory exists to test against multiple ActiveRecord versions. The supported adapter list and the SQLite 3.24.0+ / Postgres 9.5+ floors for duplicate-key updates are the constraints most likely to bite when you move a database or a framework version. The gem's own test matrix runs against the compose services, so you can reproduce it locally if you need to check a specific combination.

## Conclusion

Adopt activerecord-import when a Rails batch job spends most of its wall clock issuing one INSERT per row, especially when the rows form a parent-child graph, because the gem's recursive import is the part ActiveRecord's own insert_all does not cover. Do not adopt it if you need per-record callbacks to fire or you rely on a database adapter outside the supported list, since the README states callbacks are not run and lists the adapters explicitly. Before committing, verify two things in your own environment: that your adapter version supports the on_duplicate_key_update clause you intend to use (the README requires MySQL, SQLite 3.24.0+ or Postgres 9.5+), and that your input is column-consistent if you pass an array of hashes, because the README says an exception is raised otherwise.

## FAQ

### What does activerecord-import add to ActiveRecord?

It adds an import method (and a bulk_import alias) to ActiveRecord classes for bulk insertion. Its main feature is following associations to generate the minimal number of SQL insert statements, which the README illustrates with three statements for publishers, books and reviews.

### Which databases does activerecord-import support for on duplicate key updates?

The README states that on duplicate key updates require MySQL, SQLite 3.24.0+ or Postgres 9.5+. The repository's docker-compose.yml runs mysql:8.0 and postgis/postgis:10-2.5 for its own tests.

### What happens if I pass an array of hashes with different keys to import?

The README says an exception is raised, because using hashes only works when the columns are consistent in every hash of the array. It suggests either instantiating ActiveRecord objects first or grouping the array by keys and importing each group separately.

### Does activerecord-import run model callbacks?

The README has a callbacks section and states that callbacks are not run. Validations are a separate matter and run when the validate option is true, which is the default.

## Sources

- [Issues](https://github.com/zdennis/activerecord-import/issues)
- [License: MIT](https://github.com/zdennis/activerecord-import/blob/master/LICENSE)
- [Project website](http://www.continuousthinking.com)
- [README](https://github.com/zdennis/activerecord-import/blob/master/README.md)
- [zdennis/activerecord-import on GitHub](https://github.com/zdennis/activerecord-import)

---

Hysen Labs editorial analysis, written from the project's own repository and release notes. Cite the canonical page: https://hysenlabs.com/projects/zdennis-activerecord-import
