How to Sort a Hash by Key in Ruby (and Why You Get an Array Back)
You have a hash, you want it in key order, so you call sort and Ruby hands you an array of two element arrays. It is one of those moments where the language does exactly what it promised and still manages to surprise you.
config = { "region" => "ap-southeast-2", "adapter" => "postgresql", "pool" => 5 }
config.sort
# => [["adapter", "postgresql"], ["pool", 5], ["region", "ap-southeast-2"]]
The fix is short, but the reason is worth understanding, because it explains a whole family of Ruby behaviours and it tells you when sorting a hash is useful and when it is a waste of cycles.
Why sort returns an array
Hash does not define sort. It comes from Enumerable, which Hash includes, and every Enumerable method is built on each. Hash#each yields a two element array for every pair:
config.each { |pair| p pair }
# ["region", "ap-southeast-2"]
# ["adapter", "postgresql"]
# ["pool", 5]
So as far as Enumerable is concerned, a hash is a collection of pairs. sort, map, select, and reject all see pairs, and sort returns a sorted collection of whatever it was given. Array is the only sensible return type there, because Enumerable has no idea how to build the specific class it is iterating over.
Ruby papers over part of this itself. select, reject, and filter on a hash return hashes, because Hash overrides the Enumerable versions specifically to do that. sort, sort_by, and map are not overridden, so you get the plain Enumerable behaviour. That inconsistency is the real source of the surprise, not sort itself.
Getting a hash back
to_h converts an array of pairs back into a hash:
config.sort.to_h
# => {"adapter"=>"postgresql", "pool"=>5, "region"=>"ap-southeast-2"}
If you want to be explicit about sorting on the key rather than on the whole pair, use sort_by:
config.sort_by { |key, _value| key }.to_h
Those two are not identical. sort compares the entire [key, value] array, so when two keys are equal it falls through to comparing values. Hash keys are unique, so that tiebreak never fires and the results match. Where the difference shows up is errors: sort compares arrays, and a failed array comparison gives you a message about arrays rather than about the keys that actually caused it. sort_by { |k, _| k } fails with a message naming the key types, which is much easier to act on. I reach for sort_by by default for that reason alone.
The mixed key type trap
This is the failure people actually hit:
{ :adapter => "postgresql", "pool" => 5 }.sort
# ArgumentError: comparison of Array with Array failed
Symbols are comparable with symbols, strings are comparable with strings, and neither is comparable with the other. "pool" <=> :adapter returns nil, Array#<=> propagates that nil, and sort raises.
A hash with both symbol and string keys is usually a bug in itself, most often a Rails params hash that got partially transformed somewhere, or a merge of a config literal with a parsed YAML or JSON payload. Fixing it at the source beats working around it. If you cannot, normalize inside the sort block:
mixed.sort_by { |key, _value| key.to_s }.to_h
That coerces only the sort key, so the hash keeps its original key objects. Do not reach for transform_keys(&:to_s) here unless you actually want the keys changed, because that quietly alters the data you are about to hand to something else.
Sorting by value
The same shape works for values, which is where most real sorting happens:
counts = { "ruby" => 12, "rails" => 47, "minitest" => 3 }
counts.sort_by { |_key, value| -value }.to_h
# => {"rails"=>47, "ruby"=>12, "minitest"=>3}
Negating a number is the idiomatic descending sort and it is faster than sorting ascending and calling reverse. It only works on numerics though. For strings or anything else, sort ascending and reverse, or use sort_by { |k, v| [v, k] }.reverse.
Ruby’s sort_by is not guaranteed stable, so equal values can come back in any order, and that order can differ between Ruby versions or even between runs on large collections. If ties matter, make the sort key a tuple so there are no ties:
counts.sort_by { |key, value| [-value, key] }.to_h
Now equal counts fall back to alphabetical keys and the output is deterministic. This matters more than it sounds. A test that asserts on the top three of something with tied counts is a test that will eventually fail in CI for no reason you can reproduce locally.
When you only want the extremes, skip the sort entirely. max_by, min_by, and max_by(3) do the same job without ordering the whole collection:
counts.max_by(2) { |_key, value| value }
# => [["rails", 47], ["ruby", 12]]
Sorting nested hashes
sort is shallow. Nested hashes keep whatever order they already had, which defeats the point if you are sorting in order to produce readable or comparable output. A small recursive helper handles it:
def deep_sort(object)
case object
when Hash
object.sort_by { |key, _value| key.to_s }
.to_h { |key, value| [key, deep_sort(value)] }
when Array
object.map { |element| deep_sort(element) }
else
object
end
end
deep_sort({ b: 1, a: { d: 4, c: [{ f: 6, e: 5 }] } })
# => {:a=>{:c=>[{:e=>5, :f=>6}], :d=>4}, :b=>1}
Note that arrays are mapped, not sorted. Array order is usually meaningful, and silently reordering one would hide real differences. If you are normalizing an array of hashes for comparison, sort it explicitly on a key you choose.
When sorting is worth doing, and when it is not
Sorting does not change equality. Hash#== ignores insertion order completely:
{ a: 1, b: 2 } == { b: 2, a: 1 } # => true
So if you are sorting both sides of a test assertion to make two hashes compare equal, the sort is doing nothing and the real problem is somewhere else, usually key types, an extra key, or a value that looks equal but is not. Sorting hides nothing and fixes nothing there.
What sorting is genuinely for is anything humans or machines read positionally:
Reading output. A forty key hash printed in insertion order is much harder to scan than the same hash in alphabetical order. pp hash.sort.to_h is a cheap debugging habit.
Diffing. Two hashes built from different code paths will have different insertion orders, so a text diff of their inspect output is full of moved lines that are not changes. Sorting both sides first leaves only the real differences. This is why RubyHash sorts keys before it diffs, and it is most of why its output is readable.
Digests and cache keys. If you hash a serialized hash to build a cache key, insertion order changes the digest even though the data is identical. Sort before serializing, and deep sort if the structure is nested, or you will get cache misses that look like nothing at all.
Stable fixtures. Writing a sorted hash into a JSON or YAML fixture keeps the file’s diff small when the data changes, instead of reshuffling the whole file because a serializer changed the order of its output.
Outside those cases, sorting a hash is a pure cost: O(n log n) plus an intermediate array plus a new hash. Ruby hashes preserve insertion order for free, so if the current order is meaningful, leave it alone.
The short version
sort and sort_by come from Enumerable, they see pairs, and they return arrays. Add to_h. Prefer sort_by { |k, _| k } for clearer errors, normalize with to_s if your keys are mixed, use a tuple sort key when ties matter, and recurse if the structure is nested. Then ask whether you needed the sort at all, because for comparison you almost certainly did not.
When two hashes still will not match after you have sorted both and they still look identical on screen, the difference is a value type rather than an ordering. Try RubyHash to paste both and see exactly which key changed, including the integer that became a float and the nil that became an empty string.
Enjoyed this post?
Subscribe to get notified when we publish more Ruby and Rails content.