Rendered at 15:57:43 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
dkgs 1 days ago [-]
Hi HN!
This is Gilad, one of the creators of Pivot
We built an open-source analytics database that allows companies to run real-time analytics workloads on top of open data formats (Iceberg, Parquet), instead of needing to store data in a closed-format datastore (e.g., ClickHouse, Druid…) for real-time analytics use cases.
We've built the engine from scratch in Rust and use Arrow for in-memory data processing, alongside a thread-per-core execution model to fully control the scheduling of CPU work. To achieve the speed we wanted, we had to build a custom memory allocator, tinker with the way JOIN and GROUP BY algorithms work to fully utilize the memory bandwidth of modern CPUs, and, in general, do a lot of optimization work to ensure we achieve superior performance compared to other real-time analytics engines, while remaining committed to using only Parquet / open data formats!
We're excited to share it with the community and hope others find it as useful as we have!
niles 4 hours ago [-]
Looks interesting! Could you add a when to use / when not to use / compare with duckdb and click house section to the docs?
It's a little hard to reason about the use case for "in memory" but file based, and what the limitations are. What specific product problem did you build this to solve?
Beyond that it would be helpful to have some hardware variation in the performance claims. Does zen5 avx512 make a difference? A node with optane (memory or disk)? Threads vs clock speed? ARM vs X86? Feels like this was built with a specific cloud native deployment in mind and a certain scale and speed of storage.
dkgs 3 hours ago [-]
Hi!
Will definitely add those to the docs and landing page — really appreciate the feedback! :)
We originally built Pivot for our internal observability pipeline, where we streamed the same data into both ClickHouse (with 1 month TTL - for fast dashboards and customer-facing analytics on recent data) and Iceberg (for longer data archiving and ETL jobs). We were frustrated with this duplication (and with ClickHouse OSS being disk-based rather than S3-based!), and got curious about how fast an Iceberg-native real-time analytics database could be!
Re DuckDB — it's definitely an amazing piece of technology, but it didn't provide the full "database" experience we needed: ingestion, efficient compaction, high concurrency, cluster deployment option, and a single endpoint for our backends and dashboards to query.
Ultimately, we managed to cut the cost of our customer-facing analytics and observability pipeline by around 75%, so we figured it might be useful to others as well! :)
Re hardware variation, we definitely optimized for cloud native deployments.
While we do have some optimizations and kernels for specific instruction sets on both x86 and ARM (for example AVX-512 on x86, and with Armv8.2-A as the minimum for ARM), most of our optimizations focus on taking advantage of the memory and network bandwidth of modern hardware, both of which have increased significantly in recent years (https://pivotlake.io/docs/database/why-pivot-is-fast/#utiliz...).
This is Gilad, one of the creators of Pivot
We built an open-source analytics database that allows companies to run real-time analytics workloads on top of open data formats (Iceberg, Parquet), instead of needing to store data in a closed-format datastore (e.g., ClickHouse, Druid…) for real-time analytics use cases.
We've built the engine from scratch in Rust and use Arrow for in-memory data processing, alongside a thread-per-core execution model to fully control the scheduling of CPU work. To achieve the speed we wanted, we had to build a custom memory allocator, tinker with the way JOIN and GROUP BY algorithms work to fully utilize the memory bandwidth of modern CPUs, and, in general, do a lot of optimization work to ensure we achieve superior performance compared to other real-time analytics engines, while remaining committed to using only Parquet / open data formats!
We're excited to share it with the community and hope others find it as useful as we have!
It's a little hard to reason about the use case for "in memory" but file based, and what the limitations are. What specific product problem did you build this to solve?
Beyond that it would be helpful to have some hardware variation in the performance claims. Does zen5 avx512 make a difference? A node with optane (memory or disk)? Threads vs clock speed? ARM vs X86? Feels like this was built with a specific cloud native deployment in mind and a certain scale and speed of storage.
Will definitely add those to the docs and landing page — really appreciate the feedback! :)
We originally built Pivot for our internal observability pipeline, where we streamed the same data into both ClickHouse (with 1 month TTL - for fast dashboards and customer-facing analytics on recent data) and Iceberg (for longer data archiving and ETL jobs). We were frustrated with this duplication (and with ClickHouse OSS being disk-based rather than S3-based!), and got curious about how fast an Iceberg-native real-time analytics database could be!
Re DuckDB — it's definitely an amazing piece of technology, but it didn't provide the full "database" experience we needed: ingestion, efficient compaction, high concurrency, cluster deployment option, and a single endpoint for our backends and dashboards to query.
Ultimately, we managed to cut the cost of our customer-facing analytics and observability pipeline by around 75%, so we figured it might be useful to others as well! :)
Re hardware variation, we definitely optimized for cloud native deployments.
While we do have some optimizations and kernels for specific instruction sets on both x86 and ARM (for example AVX-512 on x86, and with Armv8.2-A as the minimum for ARM), most of our optimizations focus on taking advantage of the memory and network bandwidth of modern hardware, both of which have increased significantly in recent years (https://pivotlake.io/docs/database/why-pivot-is-fast/#utiliz...).