| ▲ | yashdotrv 8 hours ago | |||||||||||||
Hi HN, This is Yash, founding team at Parseable (https://github.com/parseablehq). We've built an open source observability data lake using Rust, that handles high-cardinality data at around 100M time series in production (https://www.parseable.com/blog/how-parseable-handles-100-mil...) Our architecture is built around columnar design, and we use Apache Arrow for in-memory columnar processing and Apache Parquet for durable columnar storage on S3-compatible object storage. In Parseable, every labels stay as columns in the data instead of becoming a large long-lived per-series index like many TSDBs. Also, one thing we’ve been thinking about a lot is how observability changes as agents become part of day-to-day engineering workflows. They're not just another service, they produce traces, tool calls, prompts, intermediate decisions, errors, costs, and sometimes sensitive business context. Observing them matters just as much as observing any other system. But it is equally important to decide where that telemetry data should reside. Our view is that teams should be able to keep these observability data close to them: in their own object storage, under their own retention, access, and compliance controls. | ||||||||||||||
| ▲ | codegeek 30 minutes ago | parent | next [-] | |||||||||||||
Your pricing page calculator is a bit strange. The minimum daily is set to 1 TB which is too high. Are you not interested in working with companies that have lesser ingestion ? | ||||||||||||||
| ▲ | goldeneye13_ 6 hours ago | parent | prev | next [-] | |||||||||||||
This looks super interesting. Question about the scale, I thought Thanos and some other Prometheus variants can handle about 100 million active time series. I would have expected your solution to scale to billions. Have you not pushed it past 100 million or am I maybe missing something. | ||||||||||||||
| ||||||||||||||
| ▲ | msandford 7 hours ago | parent | prev | next [-] | |||||||||||||
100M active time series is good information, but what's the data rate for each time series it can handle? One update per minute or 10 per second? There's a factor of 600 difference there. Neither is obviously insanely the wrong update rate. | ||||||||||||||
| ||||||||||||||
| ▲ | gustavohoa 5 hours ago | parent | prev [-] | |||||||||||||
How does the 100M active series deployment looks like? How many ingestors are there? What's each instance size? How big is the querier so that it can query across a metric with millions of active series? | ||||||||||||||
| ||||||||||||||