← Back to blog

Comparing Arrays of Hashes in Ruby Without Order Ruining It

· Lachlan Young

Comparing two hashes is usually straightforward. Comparing two arrays of hashes is where Rails tests go to die. You assert that a serializer returns the records you expect, the output looks right, the test fails anyway, and you spend twenty minutes staring at two walls of inspect output that appear to say the same thing.

Almost always the answer is one of three things: order, key type, or an extra key you did not mean to compare. Here is how to tell which, and how to write the comparison so it stops being fragile.

Array equality is positional

Array#== compares element by element, in order. Two arrays holding the same hashes in a different order are not equal, and Ruby is right about that, because arrays are ordered collections.

a = [{ id: 1, name: "Ada" }, { id: 2, name: "Grace" }]
b = [{ id: 2, name: "Grace" }, { id: 1, name: "Ada" }]

a == b  # => false

Individual hashes do not behave this way. Hash#== ignores key insertion order entirely, so {a: 1, b: 2} == {b: 2, a: 1} is true. That difference is exactly why this bug is so confusing. You have internalized that hashes do not care about order, then you wrap them in an array and suddenly order is load bearing.

The place this bites hardest is ActiveRecord. A query with no order clause gives you whatever Postgres feels like returning, which is stable enough on your machine to pass locally for months and then reorders itself in CI after a vacuum or a different row insertion order. You have not written a flaky test, you have written a test with an unspecified sort.

Fix the order, do not fight it

If the order genuinely matters to the behaviour you are testing, assert on it explicitly. Put an order on the query and keep the positional comparison, because then the ordering is part of the contract:

actual = Post.order(:published_at).map { |p| p.slice(:id, :title) }
assert_equal expected, actual

If the order does not matter, say so by sorting both sides on something deterministic before comparing:

assert_equal expected.sort_by { |h| h[:id] }, actual.sort_by { |h| h[:id] }

Sorting by a natural key is better than sorting by the whole hash, because the sort key is obvious to whoever reads the test next. When there is no single natural key, sorting by the serialized form works and is completely deterministic:

by_content = ->(h) { h.sort.to_s }

assert_equal expected.sort_by(&by_content), actual.sort_by(&by_content)

The inner h.sort matters. It sorts the hash’s pairs into a consistent order first, so two hashes with the same content but different insertion order produce the same string. Without it you are back to comparing incidental ordering.

When you truly mean “same set”

If duplicates are impossible and order is meaningless, a set comparison expresses that intent better than a sort:

require "set"

assert_equal expected.to_set, actual.to_set

Hashes work as set members because Hash#hash and Hash#eql? are defined on content, not identity, and again they ignore key insertion order. Be deliberate here though: to_set silently collapses duplicates, so a bug that returns the same record twice will pass a set comparison and fail a sorted one. Use the set form when “these are the same items” is the actual claim, and the sorted form when counts matter.

Minitest also ships assert_equal cousins that help. assert_empty(expected - actual) paired with assert_empty(actual - expected) tells you which direction the mismatch runs, which is far more useful than a single false. Array#- uses hash and eql?, so it works on hashes without any extra setup.

Narrow the comparison before you make it

Most array of hashes assertions fail on keys nobody intended to test. Timestamps are the classic case, along with database ids that change every run. Compare only what you care about:

fields = %i[title status author_id]

actual = Post.all.map { |p| p.attributes.symbolize_keys.slice(*fields) }

Two things are happening in that line, and both are common failure points on their own.

p.attributes returns string keys, always. If your expected hash uses symbols, every comparison fails and the output looks identical, because inspect renders "title" and :title almost the same way when you are scanning quickly. Either symbolize_keys the actual side or write your expectations with string keys, but pick one and be consistent across the file.

slice then drops everything you did not name, which kills the timestamp problem permanently. This is much better than the alternative of adding created_at: anything style matchers, because it makes the intent of the test obvious: these three fields, nothing else.

Finding the one element that differs

When the arrays are large, knowing they are unequal is useless. You want the index. Zip them and find the first mismatch:

expected.zip(actual).each_with_index do |(e, a), i|
  next if e == a
  puts "index #{i} differs"
  puts "  expected: #{e.inspect}"
  puts "  actual:   #{a.inspect}"
end

For collections keyed by id, comparing by key beats comparing by position, because it survives reordering and tells you about missing records too:

e_by_id = expected.index_by { |h| h[:id] }
a_by_id = actual.index_by { |h| h[:id] }

(e_by_id.keys - a_by_id.keys).each { |id| puts "missing: #{id}" }
(a_by_id.keys - e_by_id.keys).each { |id| puts "unexpected: #{id}" }

(e_by_id.keys & a_by_id.keys).each do |id|
  e, a = e_by_id[id], a_by_id[id]
  next if e == a
  changed = e.keys.reject { |k| e[k] == a[k] }
  puts "#{id} changed: #{changed.join(', ')}"
end

That is twelve lines and it turns “the arrays are not equal” into “record 42 has a different status”. index_by is ActiveSupport. In plain Ruby, each_with_object({}) or to_h { |h| [h[:id], h] } gets you the same thing.

If you are on Minitest and want this kind of output without writing it yourself, the minitest-hashdiff gem replaces the default failure output with the added, removed, and changed keys, including type changes.

The pattern worth keeping

Decide, before you write the assertion, whether order is part of what you are testing. If it is, pin it with an explicit order and compare positionally. If it is not, normalize both sides the same way: slice down to the fields under test, make the key types match, sort by a stable key, then compare. Doing the normalization in a small lambda or test helper keeps it honest, because both sides go through the same transformation and you cannot accidentally normalize only the expectation.

Almost every flaky array of hashes test is one of those two decisions left unmade.

When you have narrowed it down to a single pair of hashes that still will not match, the remaining difference is usually a value type rather than a value. Try RubyHash to paste both and see exactly which key moved, including the nil that became an empty string and the integer that became a float.

Enjoyed this post?

Subscribe to get notified when we publish more Ruby and Rails content.