Remix.run Logo
iLoveOncall 5 hours ago

So, not LLM models, right?

Also on this:

> The numbers are staggering: The DOE's light and neutron source facilities now produce tens of petabytes of data annually

Come on, petabytes are not staggering for entreprise software.

munk-a 5 hours ago | parent | next [-]

I used to work with images representing scans of brain tissue - for a full brain visualization at one horizontal slice terabytes was a common measure and the resolution of those images wasn't even particularly detailed - all the full resolution stuff was taken of tiny sub-sections of interest. This was also two decades ago - so I'm sure they've upped their game.

momojo 4 hours ago | parent [-]

My company does whole-brain scans of mice on a Zeiss Z.1. Lower resolutions are typically in the low hundreds of GB. Higher-resolutions and multichannel staining can get you in the TB range. When a typical client is doing 10's of brains it definitely adds up. But even for us (and I consider us a smaller operation) we aren't output PB's. So I'd consider the above claim still pretty impressive.

munk-a 4 hours ago | parent [-]

Yeah - our multi TB scan from two decades ago was a dolphin brain image which is a fair bit larger than mouse - and it was a stained sample that ended up being used as our demo image frequently because the contrast dyes set well and resulted in a very pretty visual overall.

I didn't meant to say that PBs of image data is common place - we had no image that approached that size - but people outside the domain of microscope scan results might be unfamiliar with just how chonky these image files would get traditionally.

hgoel 4 hours ago | parent [-]

I think a big thing that the enterprise comparison misses is that this is petabytes worth of dense data that has to be put through fairly heavy processing (some of which is currently custom for the specific experiment) and studied by a human.

It isn't just a giant database of small files and metadata blindly feeding a recommender system.

Several petabytes is definitely a staggering amount of data in that context of being analyzed by human eyes to extract some scientific value.

ben_w 4 hours ago | parent | prev [-]

I'm not sure exactly which enterprises you have in mind, but sure: quantities which can be expressed as "a year's worth fits on my desk" should not be described as "staggering", and 10 PB of hard drives will (just about) fit on my desk.

The LHC, on the other hand, that generates a petabyte a second and has to throw most of it away for obvious reasons:

https://www.itnews.com.au/news/computing-for-the-large-hadro...