I feel like this article does a depth first search on what swiss tables are, jumping head first into the tiniest implementation details, but I'm missing the breadth first search. What is the top level `struct` of a swiss table? An array of groups? Why not simplify all of it into linear open addressing, with a stride of 8 for simd? Why the triangular jumps? What problems does this design solve?
swiss tables were invented by engineers working at google's zurich office, hence the name
im surprised that go, a programming language also from google, wasn't using them!
for an excellent talk on the development of swiss tables i highly recommend this talk by Matt Kulukundis at CppCon 2017: "Designing a fast, efficient, cache-friendly hash table, step by step"
https://youtu.be/ncHmEUmJZf4
Go is much older than Swiss Tables. Since the hash table is a widely used container type and Go aspires to having a sort of "kitchen sink" stdlib I assume Go 1.0 had a hash table, and it can't be a Swiss Table because those weren't invented yet.
It's the "map" builtin. Go has a scripting-language-esque attitude of "you can build most things with arrays and hash tables". It doesn't completely preclude getting deeper but that's the general starting point.
The Rust std lib HashMap is powered by the hashbrown crate which is also a port of Swiss Tables. At a brief glance Ruby/Python don’t use this approach but I don’t see any reason why they couldn’t.
Python's hash table type, dict, promises that it remembers insertion order. If you put the mapping 5 => "dog" in first, then when we ask what is in the dict we'd get told 5 => "dog" first, duh.
Historically Python had a more conventional but very, very badly implemented hash table type, the "I can't believe it can sort"† of hash tables. Somebody wanted a hash table type which remembers insertion order because Python programmers have a bad habit of writing "Golden tests" in which that order matters even though in a good modern hash table type it's not guaranteed, so they built one. But because the built-in hash table type was garbage, this new OrderedDict type was much faster and much smaller despite solving a more difficult problem.
For a little while it was unclear if Python would decide to rewrite their dict type to have decent performance or just embrace this new better alternative and then they decided that because it's beginner friendly they will just embrace the OrderedDict and require that this type has ordering.
However, a good hash table doesn't inherently have this property, and that goes for the Swiss Table the same as other common designs. So they can't swap dict out for a Swiss Table without breaking their own promise that the dict type preserves insertion order.
If you're used to a language where this doesn't happen such as Rust, or C++ or Java or any number of other programming languages, that insertion ordering rule seems crazy, but if you've never used a programming language at all before and have never even wondered how dict works it seems obvious that this is how it should work.
ihtabs preserve iteration order and have performance competitive with swiss tables (while not requiring as high a load factor): https://github.com/vnmakarov/ihtab
Go has the same illness like rust. First var name then type. We had that back in basic and i hate it. Some argument it would be better if you declare more than one var in the same line. I never di that. Why cant we have nice things.
reply