| ▲ | jdefting 4 hours ago | |
I would be interested in seeing memory usage differences in these benchmarks. I’ve had issues with excessive memory usage in polars compared to DuckDB. I’m guessing most of the this disparity should be solved by the steaming engine. | ||
| ▲ | kaathewise 4 hours ago | parent | next [-] | |
Polars relies on threading heavily, even when streaming files. And it appears that each thread loads quite a bit of memory. I've encountered OOM issues when incrementally reading Arrow IPC files which had very large batches. Fixed it by setting $POLARS_MAX_THREADS to 1, which amusingly also improved the performance on my very narrow task. | ||
| ▲ | orlp 4 hours ago | parent | prev [-] | |
Peak memory usage per query is in the raw data (in `results/`) in the repository: https://github.com/pola-rs/polars-2.0-benchmark/. | ||