| ▲ | sanderjd 2 hours ago | ||||||||||||||||
Do you have a take on when each of these three choices is the best one? I totally agree that these are the good choices, but I still find myself hesitating about which thing to reach for when! | |||||||||||||||||
| ▲ | tomrod 2 hours ago | parent [-] | ||||||||||||||||
Depends on need. We started using PyArrow on a reporting microservice when we realized we needed no additional functionality that pandas provided since it has better data type ergonomics. DuckDb is a great go-to for SQL based transformations when working with parquet files outside a managed system like Databricks. I want to actually test duckdb versus polars with a few lower level places like iceberg on S3. | |||||||||||||||||
| |||||||||||||||||