Trending

Content tagged with "big-data"

big-data

Hacker News

Top stories from the Hacker News community• Updated less than a minute ago

132

Reddit

Top posts from tech subreddits• Updated 9 minutes ago

Hugging Face Trending

Popular models from Hugging Face• Updated 27 minutes ago

No models found

Try removing the tag filter or searching for different content.

GitHub Trending

Popular repositories from GitHub• Updated 41 minutes ago

simdjson

Parsing gigabytes of JSON per second : used by Facebook/Meta Velox, the Node.js runtime, ClickHouse, WatermelonDB, Apache Doris, Milvus, StarRocks

doris

Apache Doris is an easy-to-use, high performance and unified analytics database.

lance

Modern columnar data format for ML and LLMs implemented in Rust. Convert from parquet in 2 lines of code for 100x faster random access, vector index, and data versioning. Compatible with Pandas, DuckDB, Polars, Pyarrow, and PyTorch with more integrations coming..

starrocks

The world's fastest open query engine for sub-second analytics both on and off the data lakehouse. With the flexibility to support nearly any scenario, StarRocks provides best-in-class performance for multi-dimensional analytics, real-time analytics, and ad-hoc queries. A Linux Foundation project.

arrow

Apache Arrow is the universal columnar format and multi-language toolbox for fast data interchange and in-memory analytics

dolphinscheduler

Apache DolphinScheduler is the modern data orchestration platform. Agile to create high performance workflow with low-code

seatunnel

SeaTunnel is a multimodal, high-performance, distributed, massive data integration tool.

TDengine

High-performance, scalable time-series database designed for Industrial IoT (IIoT) scenarios

2