Remix.run Logo
bobson_dugnutt5 4 hours ago

I love polars. Did a lot of evangelizing in work to get people to give up pandas in favor of it.

anotherpaul 4 hours ago | parent | next [-]

I gave up pandas in favor of polars after someone at work did the same and I am very happy with it. Pandas API is just so much worse and much slower.

duskdozer an hour ago | parent | prev | next [-]

I guess I am a casual pandas user only. Reading a guide on migrating/differences, it's hard to see why polars would be obviously better.

bobson_dugnutt5 an hour ago | parent [-]

Here are a couple reasons:

- much faster, multithreaded by default. Read in a big csv with it and see how it feels.

- no index/MultiIndex. Pandas special treatment of index always felt like more trouble than it was worth, so no need to reset_index() everywhere.

- expressions are very portable. At first using pl.col everywhere feels like a bit much, but you can define them anywhere and then apply them to a dataframe whenever you want.

- once internalized, the syntax makes much more sense and is far more consistent compared to pandas.

Of course all depends on what your use cases are. If performance is important then I'd strongly recommend trying it out. If you just use it to have a look at the odd dataframe, maybe not worth your time as much

latexr an hour ago | parent | prev | next [-]

Taken out of context, your post looks like a conservationist who got fed up with pandas being a flagship species and made it their lifelong mission to replace them with polar bears.

This is not a criticism. As someone who doesn’t use Python, I simply found it amusing.

blitzar an hour ago | parent [-]

You should learn Boa constrictor instead of Python.

mgaunard 4 hours ago | parent | prev [-]

Both have terrible syntax that make SQL look like the most readable thing ever.

condwanaland 3 hours ago | parent | next [-]

Could not agree less. Ive always found SQL an unreadable mess but tools like polars and dplyr are such elegant ways to manipulate data.

Pandas is a mess though.

world2vec 2 hours ago | parent [-]

There's no way SQL is more unreadable than polars. IMO it's the other way around.

benrutter 2 hours ago | parent [-]

> There's no way SQL is more unreadable than polars. IMO it's the other way around.

I think on basic queries, SQL is really nice, but when stuff gets more complex, with a bunch of CTEs, let alone functions requiring loops, it becomes pretty obtuse.

geysersam an hour ago | parent | prev | next [-]

I agree sql is more elegant. The problems arise when you have to add logic on top of sql. Often I end up constructing queries via string manipulation and that is not very ergonomic. Polars api is more verbose and complex than sql but at least it's not meta-programming.

The duckdb python api is okay, but it is a bit limited, no ctes, no as of join, and it can be slow at bind/interpretation time when you do stuff like unioning multiple relations in a loop (I think that becomes O(N^2), but I might be wrong). Most issues can be worked around, but Polars is designed from the ground up to be used from python.

aquafox 3 hours ago | parent | prev | next [-]

Coming from an R/dplyr background, I agree. Compare

df.select(

  pl.col("x"),
  (pl.col("w")/pl.col("z")).alias("y")
)

with

df |> select(x, y = w/z)

orlp 3 hours ago | parent | next [-]

    from polars import col as C

    df.select(C.x, y = C.w / C.z)
dkga an hour ago | parent [-]

Still, it’s a very good approximation but still an approximation to the more ergonomic and expressive tidyverse syntax

jcattle 3 hours ago | parent | prev | next [-]

R really is/was the superior traditional data science language. Python ecosystem is slowly catching up though.

ggplot vs matplotlib

dplyr vs pandas

And I loved that everything in RStudio was so easily inspectable. Have a huge dataframe? Just look at it right in your IDE.

vovavili 2 hours ago | parent [-]

Altair and Positron should be just as good for your Polars @ Python needs. With software like Marimo notebooks and VegaFusion, Polars/Python experience starts beating R by quite a substantial margin.

bobson_dugnutt5 3 hours ago | parent | prev | next [-]

Fair point, but you can do something like

`df.select("x", y=pl.col.w/pl.col.z)`

countrymile 32 minutes ago | parent | prev [-]

Polars is a world away from pandas, but I feel that dplyr still offers the most simple and understandable introduction to data analysis for the beginner. The above is a good example of this.

bobson_dugnutt5 3 hours ago | parent | prev | next [-]

What is it about polars syntax you don't like? The fact that is very verbose? At first I wasn't a fan, but over time I've grown to really like it. That never happened to me with pandas, always felt the syntax was messy

mihaelm an hour ago | parent [-]

The verbosity takes a bit to get used too, but it sure beats the anything-goes feeling - messy as you put it - of pandas.

gpugreg 2 hours ago | parent | prev | next [-]

You can query polars data frames with SQL: https://docs.pola.rs/api/python/stable/reference/expressions...

Unfortunately, polars does not support parameterized queries, so the risk of SQL injection is extremely high.

fzumstein 3 hours ago | parent | prev [-]

I tend to agree. SQL may have been harder to write in the past (worse autocomplete than pandas/polars), but now that AI is writing the code, SQL is usually much easier to read. So DuckDB is another interesting alternative to pandas.

refactor_master 3 hours ago | parent [-]

The cool thing about polars is that you can conditionally collect expressions over many layers of business logic, and then compute the result at the end. Doing this in SQL ends up in a hodgepodge of strings and trimmed ends to please the syntax. You can also pretty effortlessly write quite complex conditionals directly in polars, and bridge it easily to the surrounding python.

I find that SQL is only easier to read with minimal abstraction, but as soon as the project gets bigger SQL becomes an unwieldy island of different that has served its purpose after we’re done with reading/writing the data.

fzumstein 3 hours ago | parent [-]

This sounds interesting! Do you have a specific example by any chance or blog post/doc references?

refactor_master 2 hours ago | parent [-]

It’s just the lazy/expression part of the API, which is really the bread and butter of polars, rather than just being “replacement syntax” for pandas. This allows you to tap into abstraction that SQL can’t keep up with:

  import polars as pl

  # 1. Base Dataset
  lazy_df = pl.LazyFrame(
    {
      "store_id": ["S01", "S02", "S03", "S04", "S05"],
      "revenue": [5000.0, 2400.0, 15000.0, 900.0, 3200.0],
      "margin": [0.45, 0.30, 0.60, 0.15, 0.50],
      "tx_count": [120, 45, 300, 20, 85],
      "returns": [5, 12, 45, 2, 8],
    }
  )

  # 2. Define Layer Abstractions
  def get_kpi_layer() -> list[pl.Expr]:
    return [
      (pl.col("returns") / pl.col("tx_count")).alias("return_rate"),
      (pl.col("revenue") / pl.col("tx_count")).alias("avg_order_value"),
    ]

  def get_threshold_layer(thresholds: dict[str, list[float]]) -> list[pl.Expr]:
    return [
      (pl.col(col) > limit).alias(f"is_{col}above{int(limit)}")
      for col, limits in thresholds.items()
      for limit in limits
    ]

  def get_interaction_layer(numeric_cols: list[str]) -> list[pl.Expr]:
    return [
      (pl.col(a) / (pl.col(b) + 1e-5)).alias(f"ratio_{a}per{b}")
      for i, a in enumerate(numeric_cols)
      for b in numeric_cols[i + 1 :]
    ]

  def get_segmentation_layer() -> list[pl.Expr]:
    return [
      pl.when(pl.col("margin") > 0.4)
      .then(pl.literal("High"))
      .otherwise(pl.literal("Low"))
      .alias("margin_profile")
    ]

  # 3. Consolidate and Execute Single Graph Pass
  thresholds = {"revenue": [1000.0, 5000.0, 10000.0], "tx_count": [50, 100, 200]}
  numeric_cols = ["revenue", "margin", "tx_count", "returns"]

  expr_pool = [
    *get_kpi_layer(),
    *get_threshold_layer(thresholds),
    *get_interaction_layer(numeric_cols),
    *get_segmentation_layer(),
  ]

  final_df = lazy_df.with_columns(expr_pool).collect()
fzumstein 14 minutes ago | parent | next [-]

awesome, thanks!

_zoltan_ an hour ago | parent | prev [-]

[flagged]