Remix.run Logo
▲ jt-s 4 hours ago

Although I find pandas a bit aggravating in many ways, for myself and my equally idiotic laboratory scientist pals, seems that it is the default way you might interface with other libraries like SciPy (i.e. they expect things as NumPy arrays or pandas dataframes). Is this a real issue or will most things happily accept a polars dataframe? We don’t work with such large datasets that speed is likely a huge concern tbh.

▲0cf8612b2e1e 3 hours ago | parent | next [-]

One nice development in this space is the narwhals library - it is a dataframe agnostic library. It allows you to seamlessly switch between pandas, polars, modlin, or any of the variations coming out.

Narwhal is still fairly new, but I expect its usage to spread since most packages only require rudimentary dataframe manipulation (set a value, math been these two columns, etc) where the limited api surface is not a problem.

Narwhals is also a much cheaper dependency to add than polars/pandas/etc so it is a somewhat easy sell to incorporate.

▲niksmather 3 hours ago | parent | prev [-]

You can convert to numpy using .to_numpy().

It's also got much better support for more complex array shapes (e.g. each row storing an array). At least it did last time I used pandas!