Polars 2.0 Defaults Every LazyFrame to the Streaming Engine. Your group_by Output Order Just Became Random.
Polars shipped a pre-release of version 2.0, and the headline change is a good one on paper: LazyFrame queries now run on the streaming engine by default instead of the in-memory engine, which the project says brings roughly a 5x speedup along with a real drop in memory usage for typical workloads. The change most people will actually get burned by isn't the speedup. It's that engine="auto" in LazyFrame.collect() now resolves to the streaming engine, and the streaming engine doesn't guarantee row order the way the in-memory engine did, on operations like group_by, join, and unpivot.
The change that throws an error, and the one that doesn't
Polars 2.0 ships five explicit breaking changes that raise an exception the moment you upgrade and run your existing code against them. Those are, in a real sense, the easy ones: your pipeline crashes, you read the traceback, you fix the call site, you move on. That's the whole failure loop working as intended.
The row-order change doesn't work that way. In Polars 1.x, engine="auto" resolved to the in-memory engine, which gave you predictable, stable row ordering after operations like group_by and join. In 2.0, that same engine="auto" call resolves to the streaming engine instead, and the streaming engine processes data in chunks without guaranteeing what order the results land in. No exception gets raised. Your code keeps running. It just returns rows in a different, effectively arbitrary order than it used to, and if anything downstream is relying on positional indexing after that operation, first row, row five, whatever, it now silently returns the wrong data.
Why this specific failure mode is worse than a crash
A crash tells you immediately that something changed. A silently reordered group_by output tells you nothing, because the code runs, the shape of the output is identical, the column names match, and every automated test that checks "does this return the right number of rows with the right columns" passes cleanly. The only thing that's wrong is which row is which, and that only shows up if a human or a downstream system happens to notice the actual values look off.
For a solo operator running a scheduled data pipeline, a nightly aggregation job, a report generator, anything that runs unattended and gets trusted by default, this is close to the worst kind of breaking change. It doesn't announce itself. It just starts producing quietly wrong output on a schedule, and the first sign is usually a stakeholder asking why last Tuesday's numbers don't match what they expected, days or weeks after the upgrade actually happened.
What to actually do before upgrading
Test against the release candidate before final 2.0 ships, specifically targeting anything that relies on row order after a group_by, join, or unpivot in a LazyFrame pipeline. If your downstream code does positional indexing, df[0], .head(1), anything that assumes "the first row is the one I expect", audit those call sites now rather than after the upgrade. Where row order actually matters for correctness, not just habit, add an explicit .sort() after the operation rather than relying on incidental ordering from whichever engine happens to run. That's arguably the more correct pattern regardless of this change, since relying on unspecified ordering was always a soft assumption, Polars 2.0 just stops letting you get away with it for free.
The honest take
The streaming-by-default change itself is the right call, and a roughly 5x performance improvement for typical workloads is a real, substantial win that most Polars users will benefit from without doing anything. My complaint isn't with the direction, it's with how a change of this consequence ships without an exception path. A silent behavior change to something as fundamental as row ordering deserves at minimum an opt-in deprecation warning during a transition period, even if the long-term default is the right one. If you run any unattended Polars pipeline, don't wait for the final 2.0 release to find out whether this affects you.
Author
Lukas
@lukcombinator