Remix.run Logo
▲ desipenguin 7 hours ago

From recent Python Bytes podcast (https://pythonbytes.fm/episodes/show/496/a-lake-house-in-sea...)

> 1 Billion Row Challenge benchmark: Pandas took 4m28s vs. Polars 5.04s and DuckDB 5.19s — DuckDB also used 19x less memory

Python Vs Rust : In terms for speed - No comparison

(The above episode transcript has a link to blog post titled "Pandas should go extinct" )

▲jszymborski 2 hours ago | parent [-]

I've nearly entirely switched to DuckDB for anything more than like 500 or 1,000 rows or if there are a tonne of columns.

Polars is great, but I'm just too used to the Pandas API to use it as a replacement for the cases where DuckDB is overkill.

▲entropicdrifter an hour ago | parent [-]

That's a shame, because Pandas has a really quirky/legacy-burdened API and Polars is super clean by comparison. As someone who had Spark and Pandas experience before switching, Polars felt like Pyspark without the added mental overhead of needing you to think about multi-worker-node parallelism