Data-integratie

240 projecten in Data & analyse

Toont top 200 van 240 op Atlas Score; verfijn met de filters hierboven.

ShardingSphere

@apache

Empowering Data Intelligence with Distributed SQL for Sharding, Scalability, and Security Across All Databases.

Apache-2.0 Commit 1 dag geleden ★ 20.805
80 5/5 gemeten

Kestra

@kestra-io

Event-driven, language-agnostic platform to create, schedule, and monitor workflows. In code. Coordinate data pipelines and tasks such as ETL and ELT.

Zelf te hosten Commit 3 dagen geleden ★ 28.386
76 5/5 gemeten

Apache Airflow

@apache

Platform to programmatically author, schedule, and monitor workflows.

Zelf te hosten Commit 1 dag geleden ★ 46.994
76 5/5 gemeten

Vector

@vectordotdev

A high-performance observability data pipeline.

MPL-2.0 Commit 1 dag geleden ★ 22.626
75 4/5 gemeten

multiwoven

@Multiwoven

🔥🔥🔥 Open source Reverse ETL - alternative to hightouch and census.

Zelf te hosten Commit 2 dagen geleden ★ 1.677
74 4/5 gemeten

Prefect

@PrefectHQ

Prefect is a workflow orchestration framework for building resilient data pipelines in Python.

Apache-2.0 Commit 1 dag geleden ★ 23.935
72 5/5 gemeten

whodb

@clidey

Where data access meets operational intelligence

Zelf te hosten Commit 1 dag geleden ★ 5.031
72 4/5 gemeten

cloudquery

@cloudquery

Data pipelines for cloud config and security data. Build cloud asset inventory, CSPM, FinOps, and vulnerability management solutions. Extract from AWS, Azure, GCP, and 70+ cloud and SaaS sources.

MPL-2.0 Commit 2 dagen geleden ★ 6.528
71 5/5 gemeten

RudderStack

@rudderlabs

Collect, unify, transform, and store your customer data, and route it to a wide range of common, popular marketing, sales, and product tools (alternative to Segment).

Zelf te hosten Commit 1 dag geleden ★ 4.490
70 5/5 gemeten

mage-ai

@mage-ai

🧙 Build, run, and manage data pipelines for integrating and transforming data.

Apache-2.0 Commit 17 dagen geleden ★ 8.826
70 4/5 gemeten

sqlmesh

@SQLMesh

Scalable and efficient data transformation framework - backwards compatible with dbt.

Apache-2.0 Commit 3 dagen geleden ★ 3.297
70 4/5 gemeten

taipy

@Avaiga

Turns Data and AI algorithms into production-ready web applications in no time.

Apache-2.0 Commit 2 maanden geleden ★ 19.434
69 4/5 gemeten

duckle

@slothflowlabs

Open-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.

Zelf te hosten Commit 3 dagen geleden ★ 1.330
68 4/5 gemeten

flink-cdc

@apache

Flink CDC is a streaming data integration tool

Apache-2.0 Commit 2 dagen geleden ★ 6.480
68 4/5 gemeten

myduckserver

@apecloud

Unified MySQL, Postgres & FlightSQL Server, Powered by DuckDB.

Apache-2.0 Commit 9 dagen geleden ★ 592
68 4/5 gemeten

reductstore

@reductstore

High-performance, time-indexed object storage for robotics and industrial IoT

Maintainer in de EU Commit 1 dag geleden ★ 372
67 4/5 gemeten

desbordante-core

@Desbordante

Desbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.

AGPL-3.0 Commit 1 dag geleden ★ 509
67 4/5 gemeten

devlake

@apache

Apache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.

Apache-2.0 Commit 1 dag geleden ★ 3.149
67 4/5 gemeten

cocoindex

@cocoindex-io

Incremental engine for long horizon agents 🌟 Star if you like it!

Apache-2.0 Commit 4 dagen geleden ★ 11.607
67 4/5 gemeten

risingwave

@risingwavelabs

Event streaming platform for agentic AI. Continuously ingest, transform, and serve event streams in real time, at scale.

Apache-2.0 Commit 1 dag geleden ★ 9.350
67 4/5 gemeten

hudi

@apache

Upserts, Deletes And Incremental Processing on Big Data.

Apache-2.0 Commit 1 dag geleden ★ 6.274
67 5/5 gemeten

data-engineering-wiki

@data-engineering-community

The best place to learn data engineering. Built and maintained by the data engineering community.

CC0-1.0 Commit 13 dagen geleden ★ 2.030
67 4/5 gemeten

squid-sdk

@subsquid

TypeScript ETL toolkit for indexing Ethereum, Solana, and Substrate data, sourced from SQD Network.

Apache-2.0 Commit 3 dagen geleden ★ 1.341
67 4/5 gemeten

airbyte

@airbytehq

Open-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.

Zelf te hosten Commit 1 dag geleden ★ 22.142
66 5/5 gemeten

hamilton

@apache

Apache Hamilton helps data scientists and engineers define testable, modular, self-documenting dataflows, that encode lineage/tracing and metadata. Runs and scales everywhere python does.

Apache-2.0 Commit 2 dagen geleden ★ 2.599
66 4/5 gemeten

Daft

@Eventual-Inc

High-performance data engine for AI and multimodal workloads. Process images, audio, video, and structured data at any scale

Apache-2.0 Commit 2 dagen geleden ★ 5.785
66 4/5 gemeten

mlops-python-package

@fmind

A comprehensive Python package template to kickstart and standardize your MLOps initiatives and data pipelines.

Maintainer in de EU Commit 7 dagen geleden ★ 1.418
66 4/5 gemeten

dbt

@dbt-labs

dbt enables data analysts and engineers to transform their data using the same practices that software engineers use to build applications.

Apache-2.0 Commit 1 dag geleden ★ 13.936
66 5/5 gemeten

debezium

@debezium

Change data capture for a variety of databases. Please log issues at https://github.com/debezium/dbz/issues.

Apache-2.0 Commit 1 dag geleden ★ 13.160
66 5/5 gemeten

jitsu

@jitsucom

Jitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for modern data teams. Set-up a real-time data pipeline in minutes, not days

MIT Commit 3 dagen geleden ★ 5.093
66 5/5 gemeten

meltano

@meltano

Meltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.

MIT Commit 3 dagen geleden ★ 2.637
66 4/5 gemeten

olake

@datazip-inc

OLake - Fastest Databases, Kafka & S3 Replication to Apache Iceberg with Table optimization (Called OLake Fusion). ⚡ Efficient, quick and scalable data ingestion for real-time analytics. Supported sources : Postgres, MongoDB, MySQL, Oracle, MSSql, DB2, Kafka, S3.

Apache-2.0 Commit 1 dag geleden ★ 1.462
66 4/5 gemeten

pgsync

@toluaina

Postgres, MySQL, or MariaDB to Elasticsearch/OpenSearch sync

MIT Commit 7 dagen geleden ★ 1.423
66 4/5 gemeten

egeria

@odpi

Egeria core

Apache-2.0 Commit 1 dag geleden ★ 924
66 5/5 gemeten

DataFlow-Engine

@risesoft-y9

数据流引擎是一款面向数据集成、数据同步、数据交换、数据共享、任务配置、任务调度的底层数据驱动引擎。数据流引擎采用管执分离、多流层、插件库等体系应对大规模数据任务、数据高频上报、数据高频采集、异构数据兼容的实际数据问题。

GPL-3.0 Commit 12 dagen geleden ★ 697
66 4/5 gemeten

maestro

@Netflix

Maestro: Netflix’s Workflow Orchestrator

Apache-2.0 Commit 6 dagen geleden ★ 3.843
65 4/5 gemeten

ape-dts

@apecloud

ApeCloud's Data Transfer Suite, written in Rust. Provides ultra-fast data replication between MySQL, PostgreSQL, Redis, MongoDB, Kafka and ClickHouse, ideal for disaster recovery (DR) and migration scenarios.

AGPL-3.0 Commit 6 dagen geleden ★ 603
65 4/5 gemeten

aws-sdk-pandas

@aws

pandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).

Apache-2.0 Commit 2 dagen geleden ★ 4.119
65 4/5 gemeten

hop

@apache

Hop Orchestration Platform

Apache-2.0 Commit 1 dag geleden ★ 1.478
65 4/5 gemeten

ReplicaDB

@osalvador

ReplicaDB is open source tool for database replication, designed for efficiently transferring bulk data between relational and non-relational databases

Maintainer in de EU Commit 5 dagen geleden ★ 499
65 4/5 gemeten

od

@kokes

Česká otevřená data

Maintainer in de EU Commit 1 maand geleden ★ 138
64 4/5 gemeten

apple-notes-exporter

@kzaremski

MacOS app written in Swift that bulk exports Apple Notes (including iCloud Notes) to a multitude of formats preserving note folder structure.

GPL-3.0 Commit 9 dagen geleden ★ 722
64 4/5 gemeten

pudl

@catalyst-cooperative

The Public Utility Data Liberation Project provides analysis-ready energy system data to climate advocates, researchers, policymakers, and journalists.

MIT Commit 1 dag geleden ★ 608
63 5/5 gemeten

odd-platform

@opendatadiscovery

First open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.

Apache-2.0 Commit 6 dagen geleden ★ 1.433
63 4/5 gemeten

dagster

@dagster-io

An orchestration platform for the development, production, and observation of data assets.

Apache-2.0 Commit 3 dagen geleden ★ 16.207
63 5/5 gemeten

Addax

@wgzhao

Actively maintained successor to Alibaba DataX — a fast, versatile, open-source ETL tool for 20+ RDBMS and NoSQL data sources.

Apache-2.0 Commit 3 dagen geleden ★ 1.442
63 4/5 gemeten

datacontract-cli

@datacontract

Enforce Data Contracts

MIT Commit 3 dagen geleden ★ 1.072
63 4/5 gemeten

gspread-pandas

@Paradigmllc

Read and write Google Sheets as pandas DataFrames — column-matched appends, real dtypes, and layout detection for sheets that don't start at A1.

BSD-3-Clause Commit 2 maanden geleden ★ 414
62 4/5 gemeten

spark-excel

@nightscape

A Spark plugin for reading and writing Excel files

Maintainer in de EU Commit 4 dagen geleden ★ 524
62 4/5 gemeten

tis

@datavane

Support agile Ontology DataOps Based on Flink, DataX and Flink-CDC with Web-UI

Apache-2.0 Commit 11 dagen geleden ★ 1.416
62 4/5 gemeten

open-data-contract-standard

@bitol-io

Home of the Open Data Contract Standard (ODCS).

Apache-2.0 Commit 17 dagen geleden ★ 1.152
62 4/5 gemeten

rocky

@rocky-data

A SQL transformation engine that type-checks your whole pipeline and catches breaking changes before they run — branches, replay, column-level lineage, compile-time contracts, per-model cost. Adapters: Databricks, Snowflake, BigQuery, DuckDB. Single static Rust binary. Apache 2.0.

Maintainer in de EU Commit 1 dag geleden ★ 301
62 4/5 gemeten

riko

@nerevu

A Python stream processing engine modeled after Yahoo! Pipes

MIT Commit 3 dagen geleden ★ 1.607
61 5/5 gemeten

monopoly

@benjamin-awd

Monopoly is a Python library & CLI that converts bank statement PDFs to CSV.

AGPL-3.0 Commit 2 dagen geleden ★ 426
61 4/5 gemeten

pyjanitor

@pyjanitor-devs

Clean APIs for data cleaning. Python implementation of R package Janitor

MIT Commit 1 dag geleden ★ 1.500
61 5/5 gemeten

conduit

@ConduitIO

Conduit streams data between data stores. Kafka Connect replacement. No JVM required.

Apache-2.0 Commit 2 dagen geleden ★ 611
61 4/5 gemeten

filesql

@nao1215

loads CSV, TSV, LTSV, JSON, JSONL, Parquet, XLSX, ACH, and Fedwire files into SQLite; includes prep and frame for cleanup and in-memory transforms

MIT Commit 2 dagen geleden ★ 386
60 4/5 gemeten

sling-cli

@slingdata-io

Sling is a CLI tool that extracts data from a source storage/database and loads it in a target storage/database.

GPL-3.0 Commit vandaag ★ 910
60 5/5 gemeten

koheesio

@Nike-Inc

Python framework for building efficient data pipelines. It promotes modularity and collaboration, enabling the creation of complex pipelines from simple, reusable components.

Apache-2.0 Commit 27 dagen geleden ★ 818
60 4/5 gemeten

dbt-databricks

@databricks

A dbt adapter for Databricks.

Apache-2.0 Commit 3 dagen geleden ★ 379
60 4/5 gemeten

spark

@dataflint

Drop-in replacement for Apache Spark UI

Apache-2.0 Commit 1 maand geleden ★ 491
59 4/5 gemeten

chunjun

@DTStack

A data integration framework

Apache-2.0 Commit 10 maanden geleden ★ 4.100
59 4/5 gemeten

soda-core

@sodadata

Data Contracts engine for the modern data stack. https://www.soda.io

Maintainer in de EU Commit 4 dagen geleden ★ 2.430
59 4/5 gemeten

zerocode

@authorjapps

zerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Databases and Load scenarios. all defined in simple JSON or YAML — with zero coding.

Apache-2.0 Commit 6 dagen geleden ★ 1.014
59 5/5 gemeten

datavines

@datavane

Know your data better!Datavines is Next-gen Data Observability Platform, support metadata manage and data quality.

Apache-2.0 Commit 2 maanden geleden ★ 765
59 4/5 gemeten

bacalhau

@bacalhau-project

Community-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.

Apache-2.0 Commit 1 dag geleden ★ 872
58 5/5 gemeten

extractor

@lightfeed

Use LLMs to robustly extract web data

Apache-2.0 Commit 4 maanden geleden ★ 321
58 4/5 gemeten

CommonCoreOntologies

@CommonCoreOntology

The Common Core Ontology Repository holds the current released version of the Common Core Ontology suite.

BSD-3-Clause Commit 3 dagen geleden ★ 381
58 4/5 gemeten

snowpark-python

@snowflakedb

Snowflake Snowpark Python API

Apache-2.0 Commit 1 dag geleden ★ 341
58 4/5 gemeten

dbt-trino

@starburstdata

The Trino (https://trino.io/) adapter plugin for dbt (https://getdbt.com)

Maintainer in de EU Commit 5 dagen geleden ★ 267
58 5/5 gemeten

dbt-sqlserver

@dbt-msft

dbt adapter for SQL Server and Azure SQL

MIT Commit 1 dag geleden ★ 258
58 4/5 gemeten

lakeFS

@treeverse

lakeFS - Data version control for your data lake | Git for data

Commit 3 dagen geleden ★ 5.542
58 5/5 gemeten

fluvio

@fluvio-community

🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.

Apache-2.0 Commit 29 dagen geleden ★ 5.259
58 4/5 gemeten

recce

@DataRecce

The data-validation toolkit for enhanced dbt (data build tool) PR review

Apache-2.0 Commit 1 dag geleden ★ 480
58 4/5 gemeten

DataMate

@ModelEngine-Group

DataMate is an enterprise-level data processing platform designed for model fine-tuning and RAG retrieval.

MIT Commit 1 maand geleden ★ 368
58 4/5 gemeten

Flowfile

@Edwardvaneechoud

Flowfile is a visual ETL tool and Python library combining drag-and-drop workflows with Polars dataframes. Build data pipelines visually, define flows programmatically with a Polars-like API, and export to standalone Python code. Perfect for fast, intuitive data processing from development to production.

MIT Commit 1 dag geleden ★ 363
58 4/5 gemeten

kafka-connect-file-pulse

@streamthoughts

🔗 A multipurpose Kafka Connect connector that makes it easy to parse, transform and stream any file, in any format, into Apache Kafka

Maintainer in de EU Commit 3 maanden geleden ★ 350
58 5/5 gemeten

DataEngineeringPilipinas

@ogbinar

Data Engineering Pilipinas is a community for data engineers, data analysts, data scientists, developers, AI / ML engineers, and users of closed and open source data tools and methods / techniques in the Philippines. Data Engineering Pilipinas is a PyData group.

MIT Commit 7 dagen geleden ★ 237
58 4/5 gemeten

dflib

@dflib

In-memory Java DataFrame library

Apache-2.0 Commit 3 dagen geleden ★ 325
57 4/5 gemeten

pathway

@pathwaycom

Python ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.

Commit 3 dagen geleden ★ 62.220
57 4/5 gemeten

dbt-coves

@datacoves

CLI tool for dbt users to simplify creation of staging models (yml and sql) files

Apache-2.0 Commit 7 dagen geleden ★ 282
57 4/5 gemeten

nomenklatura

@opensanctions

Framework and command-line tools for integrating FollowTheMoney data streams from multiple sources

Maintainer in de EU Commit 3 dagen geleden ★ 267
57 5/5 gemeten

CloudTAK

@dfpc-coe

TAK Compatible, browser based Common Operation Picture & Situational Awareness tool

AGPL-3.0 Commit 1 dag geleden ★ 134
56 4/5 gemeten

qsv

@dathere

Blazing-fast Data-Wrangling toolkit

Commit vandaag ★ 3.797
56 4/5 gemeten

zdh_web

@zhaoyachao

大数据采集,抽取平台,zdh_web是zdh系列服务的可视化管理平台,包含数据采集,调度,权限,审批流,私域营销等模块

Apache-2.0 Commit 4 maanden geleden ★ 536
56 4/5 gemeten

growthbook

@growthbook

Open Source Feature Flags, Experimentation, and Product Analytics

Commit 1 dag geleden ★ 8.445
56 4/5 gemeten

wexflow

@aelassas

Workflow Automation Engine

MIT Commit 8 dagen geleden ★ 842
56 5/5 gemeten

automate-dv

@Datavault-UK

A free to use dbt package for creating and loading Data Vault 2.0 compliant Data Warehouses (powered by dbt, an open source data engineering tool, registered trademark of dbt Labs)

Apache-2.0 Commit 8 maanden geleden ★ 596
56 4/5 gemeten

lakehouse-engine

@adidas

The Lakehouse Engine is a configuration driven Spark framework, written in Python, serving as a scalable and distributed engine for several lakehouse algorithms, data flows and utilities for Data Products.

Apache-2.0 Commit 1 maand geleden ★ 293
56 4/5 gemeten

morph-kgc

@morph-kgc

Powerful RDF knowledge graph generation with RML mappings

Maintainer in de EU Commit 3 maanden geleden ★ 291
56 4/5 gemeten

hflow

@Hebbian-Robotics

SDK for robotics teams to verify the quality of their data used for AI model training.

Apache-2.0 Commit 1 dag geleden ★ 283
56 4/5 gemeten

steampipe-plugin-aws

@turbot

Use SQL to instantly query AWS resources across regions and accounts. Open source CLI. No DB required.

Apache-2.0 Commit 2 dagen geleden ★ 231
56 5/5 gemeten

starflow

@starlake-ai

Declarative text based tool for data analysts and engineers to extract, load, transform and orchestrate their data pipelines.

Maintainer in de EU Commit 1 dag geleden ★ 213
56 4/5 gemeten

skills

@dagster-io

A collection of AI skills for working with Dagster

Apache-2.0 Commit 7 dagen geleden ★ 208
55 4/5 gemeten

harmonypy

@slowkow

🎼 Integrate multiple high-dimensional datasets with fuzzy k-means and locally linear adjustments.

GPL-3.0 Commit 12 dagen geleden ★ 284
55 5/5 gemeten

mloda

@mloda-ai

mloda.ai - Open Data Access for AI and ML. Plugin-based. Traceable. Framework-agnostic.

Maintainer in de EU Commit 1 dag geleden ★ 92
55 4/5 gemeten

streamable

@ebonnal

sync/async iterable streams for Python

Apache-2.0 Commit 1 dag geleden ★ 332
55 4/5 gemeten

connect

@redpanda-data

Fancy stream processing made operationally mundane

Commit 1 dag geleden ★ 8.769
54 4/5 gemeten

awesome-engineering-articles

@ashishps1

A curated collection of 300+ engineering blog articles from top tech companies. Learn how the best engineering teams solve real-world problems at scale.

MIT Commit 7 maanden geleden ★ 1.425
54 4/5 gemeten

sqrl

@DataSQRL

Agentic Data Engineering Harness for building data pipelines, data products, data APIs, and data lakes autonomously

Apache-2.0 Commit 2 dagen geleden ★ 228
54 4/5 gemeten

xonsh

@xonsh

🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.

Commit 1 dag geleden ★ 9.658
53 5/5 gemeten

superglue

@superglue-ai

superglue (YC W25) builds integrations and tools from natural language. Get production-grade tools for long tail and enterprise systems.

Commit 1 maand geleden ★ 2.063
53 4/5 gemeten

zer0share

@zer0quant

A 股、期货、期权数据本地化管道:Tushare Pro 拉取 → Parquet 分区存储 → DuckDB 本地查询,支持增量同步与定时调度

MIT Commit 1 dag geleden ★ 185
53 4/5 gemeten

ingestr

@bruin-data

ingestr is a CLI tool to copy data between any databases with a single command seamlessly.

Commit 3 dagen geleden ★ 3.982
53 4/5 gemeten

usaspending-api

@fedspendingtransparency

Server application to serve U.S. federal spending data via a RESTful API

CC0-1.0 Commit 3 dagen geleden ★ 467
53 4/5 gemeten

mcp-cn-commerce

@TonyWang-hub

中国电商 MCP 连接器:8 个平台、155 个已注册工具。开源标准版,含能力矩阵与实际工程验收记录;真实商家联调状态逐项公开。Chinese commerce MCP connectors with documented capabilities and validation status.

MIT Commit 7 dagen geleden ★ 67
52 4/5 gemeten

datajoint-python

@datajoint

Relational Workflows: where database schemas define executable data pipelines.

Apache-2.0 Commit 19 dagen geleden ★ 197
52 4/5 gemeten

snowplow

@snowplow

The leader in Customer Data Infrastructure

Apache-2.0 Commit 3 maanden geleden ★ 7.034
52 5/5 gemeten

quilt

@quiltdata

Quilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.

Apache-2.0 Commit 1 dag geleden ★ 1.370
52 5/5 gemeten

beginner_de_project

@josephmachado

Beginner data engineering project - batch edition

MIT Commit 6 maanden geleden ★ 595
52 4/5 gemeten

databricks_bootcamp_2026

@DataWithBaraa

End-to-end Data Lakehouse project built on Databricks, following the Medallion Architecture (Bronze, Silver, Gold). Covers real-world data engineering and analytics workflows using Spark, PySpark, SQL, Delta Lake, and Unity Catalog. Designed for learning, portfolio building, and job interviews.

MIT Commit 8 maanden geleden ★ 425
52 4/5 gemeten

wingfoil

@wingfoil-io

ultra low latency graph based stream processing framework

Apache-2.0 Commit 1 dag geleden ★ 224
52 4/5 gemeten

paperetl

@neuml

📄 ⚙️ ETL processes for medical and scientific papers

Apache-2.0 Commit 10 maanden geleden ★ 697
51 5/5 gemeten

cnpj-data-pipeline

@caiopizzol

Pipeline open-source que baixa e processa os dados da Receita Federal para PostgreSQL

Commit 20 dagen geleden ★ 309
51 4/5 gemeten

pentaho-kettle

@pentaho

Pentaho Data Integration ( ETL ) a.k.a Kettle

Commit 1 dag geleden ★ 8.399
51 4/5 gemeten

sql-translator

@whoiskatrin

SQL Translator is a tool for converting natural language queries into SQL code using artificial intelligence. This project is 100% free and open source.

MIT Commit 1 jaar geleden ★ 4.328
50 4/5 gemeten

amphi-etl

@amphi-ai

visual data prep powered by python

Commit 1 maand geleden ★ 1.410
50 4/5 gemeten

ChoETL

@Cinchoo

ETL framework for .NET (Parser / Writer for CSV, Flat, Xml, JSON, Key-Value, Parquet, Yaml, Avro formatted files)

MIT Commit 3 maanden geleden ★ 860
50 5/5 gemeten

extract

@ICIJ

A cross-platform command line tool for parallelised content extraction and analysis.

MIT Commit 28 dagen geleden ★ 259
50 5/5 gemeten

ethereum-etl

@blockchain-etl

Python scripts for ETL (extract, transform and load) jobs for Ethereum blocks, transactions, ERC20 / ERC721 tokens, transfers, receipts, logs, contracts, internal transactions. Data is available in Google BigQuery https://goo.gl/oY5BCQ

MIT Commit 8 maanden geleden ★ 3.130
50 5/5 gemeten

flow

@estuary

🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊

Commit 1 dag geleden ★ 980
50 5/5 gemeten

recap

@gabledata

Work with your web service, database, and streaming schemas in a single format.

MIT Commit 9 maanden geleden ★ 352
50 4/5 gemeten

public-datasets-pipelines

@GoogleCloudPlatform

Cloud-native, data onboarding architecture for Google Cloud Datasets

Apache-2.0 Commit 3 maanden geleden ★ 180
49 5/5 gemeten

data-science-on-gcp

@GoogleCloudPlatform

Source code accompanying book: Data Science on the Google Cloud Platform, Valliappa Lakshmanan, O'Reilly 2017

Apache-2.0 Commit 7 maanden geleden ★ 1.431
49 5/5 gemeten

Data-engineering-nanodegree

@Flor91

Projects done in the Data Engineering Nanodegree by Udacity.com

MIT Commit 7 maanden geleden ★ 273
49 4/5 gemeten

go-etl

@Breeze0806

go-etl is a toolset for data extraction, transformation and loading.

Apache-2.0 Commit 4 maanden geleden ★ 192
49 5/5 gemeten

incubator-graphar

@apache

An open source, standard data file format for graph data storage and retrieval.

Apache-2.0 Commit 12 dagen geleden ★ 374
49 4/5 gemeten

opendbt

@memiiso

Make dbt great again! Extend dbt with plugins, local docs and custom adapters — fast, safe, and developer-friendly

Apache-2.0 Commit 2 maanden geleden ★ 295
49 4/5 gemeten

flowcraft

@gorango

A lightweight workflow engine

MIT Commit 3 maanden geleden ★ 210
49 4/5 gemeten

pg_background

@vibhorkum

Production-grade PostgreSQL extension to execute arbitrary SQL in background worker processes — with async execution, autonomous transactions, cookie-protected handles, cancellation, progress reporting, and observability.

Commit 1 maand geleden ★ 257
48 5/5 gemeten

Exchangis

@WeBankFinTech

Exchangis is a lightweight,highly extensible data exchange platform that supports data transmission between structured and unstructured heterogeneous data sources

Apache-2.0 Commit 11 maanden geleden ★ 461
48 4/5 gemeten

databricks-code-practice

@jrlasak

Practice Databricks coding skills with hands-on exercises. Import into Databricks Free Edition, write code, run assertions, check pass/fail. Covers Delta Lake, Spark SQL, PySpark, Auto Loader, medallion architecture, window functions, and more.

Maintainer in de EU Commit 2 maanden geleden ★ 278
48 4/5 gemeten

pypi-duck-flow

@mehd-io

end-to-end data engineering project to get insights from PyPi using python, duckdb, MotherDuck

Maintainer in de EU Commit 13 dagen geleden ★ 238
48 4/5 gemeten

DataCleaner

@datacleaner

The premier open source Data Quality solution

LGPL-3.0 Commit 2 maanden geleden ★ 651
47 5/5 gemeten

redun

@insitro

Yet another redundant workflow engine

Apache-2.0 Commit 2 maanden geleden ★ 603
47 5/5 gemeten

bulk-writer

@jbogard

Provides guidance for fast ETL jobs, an IDataReader implementation for SqlBulkCopy (or the MySql or Oracle equivalents) that wraps an IEnumerable, and libraries for mapping entites to table columns.

MIT Commit 10 maanden geleden ★ 242
47 4/5 gemeten

pglogical

@2ndQuadrant

Logical Replication extension for PostgreSQL 17, 16, 15, 14, 13, 12, 11, 10, 9.6, 9.5, 9.4 (Postgres), providing much faster replication than Slony, Bucardo or Londiste, as well as cross-version upgrades.

Commit 2 maanden geleden ★ 1.239
47 5/5 gemeten

qData

@qiantongtech

qData is an open-source data governance and data development platform that integrates ETL, data development, metadata management, data quality, data assets, API services, and AI-powered data Q&A.

Commit 4 dagen geleden ★ 609
47 4/5 gemeten

gusty

@pipeline-tools

Making DAG construction easier

MIT Commit 2 maanden geleden ★ 286
47 4/5 gemeten

airflow-dbt-python

@tomasfarias

A collection of Airflow operators, hooks, and utilities to elevate dbt to a first-class citizen of Airflow.

MIT Commit 11 dagen geleden ★ 215
47 5/5 gemeten

complete-dbt-bootcamp-zero-to-hero

@zoltanctoth

Supplementary Materials for the The Complete dbt (Data Build Tool) Bootcamp Udemy course

Maintainer in de EU Commit 5 dagen geleden ★ 832
46 4/5 gemeten

tributary

@1kbgz

Streaming reactive and dataflow graphs in Python

Apache-2.0 Commit 3 maanden geleden ★ 466
46 4/5 gemeten

radient

@fzliu

Radient turns many data types (not just text) into vectors for similarity search, RAG, regression analysis, and more.

BSD-2-Clause Commit 7 maanden geleden ★ 281
46 4/5 gemeten

go-streams

@reugn

A lightweight stream processing library for Go

MIT Commit 9 maanden geleden ★ 2.173
46 5/5 gemeten

seatunnel-web

@apache

SeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).

Apache-2.0 Commit 8 maanden geleden ★ 883
46 4/5 gemeten

metorikku

@YotpoLtd

A simplified, lightweight ETL Framework based on Apache Spark

MIT Commit 27 dagen geleden ★ 589
46 5/5 gemeten

meilisync

@long2ice

Realtime sync data from MySQL/PostgreSQL/MongoDB to Meilisearch

Apache-2.0 Commit 4 maanden geleden ★ 384
46 4/5 gemeten

nichenetr

@saeyslab

NicheNet: predict active ligand-target links between interacting cells

Maintainer in de EU Commit 3 dagen geleden ★ 680
45 4/5 gemeten

practical-data-engineering

@ssp-data

Practical Data Engineering: A Hands-On Real-Estate Project Guide

Commit 3 maanden geleden ★ 825
45 4/5 gemeten

dagster-open-platform

@dagster-io

Dagster Labs' open-source data platform, built with Dagster.

Commit 17 dagen geleden ★ 475
45 4/5 gemeten

modern-polars

@kevinheavey

Code and data for the Modern Polars book

Maintainer in de EU Commit 3 maanden geleden ★ 234
45 4/5 gemeten

fhir-data-pipes

@ohs-foundation

A collection of tools for extracting FHIR resources and analytics services on top of that data.

Apache-2.0 Commit 6 dagen geleden ★ 225
45 4/5 gemeten

metl

@jumpmindinc

Metl is a simple, web-based integration platform that allows for several different styles of data integration including messaging, file based Extract/Transform/Load (ETL), and remote procedure invocation via Web Services. Read more at www.jumpmind.com/products/metl/overview

GPL-3.0 Commit 3 maanden geleden ★ 215
44 4/5 gemeten

jaffle-shop

@dbt-labs

🥪🦘 An open source sandbox project exploring dbt workflows via a fictional sandwich shop's data.

Commit 12 dagen geleden ★ 366
44 4/5 gemeten

every-single-day-i-tldr

@sderosiaux

A daily digest of the articles or videos I've found interesting, that I want to share with you.

Commit 26 dagen geleden ★ 328
43 4/5 gemeten

dcs-core

@datachecks

Open Source Data Quality Monitoring.

Apache-2.0 Commit 8 maanden geleden ★ 179
43 4/5 gemeten

sql-ultimate-course

@DataWithBaraa

The most comprehensive SQL guide from a real-world expert! Learn everything from basics to advanced queries, optimizations, and real-world SQL

MIT Commit 1 jaar geleden ★ 1.439
43 4/5 gemeten

PyAirbyte

@airbytehq

PyAirbyte brings the power of Airbyte to every Python developer. Powers the Airbyte Cloud Replication MCP.

Commit 2 dagen geleden ★ 343
43 4/5 gemeten

kuwala

@kuwala-io

Kuwala is the no-code data platform for BI analysts and engineers enabling you to build powerful analytics workflows. We are set out to bring state-of-the-art data engineering tools you love, such as Airbyte, dbt, or Great Expectations together in one intuitive interface built with React Flow. In addition we provide third-party data into data science models and products with a focus on geospatial data. Currently, the following data connectors are available worldwide: a) High-resolution demographics data b) Point of Interests from Open Street Map c) Google Popular Times

Maintainer in de EU Commit 4 jaaren geleden ★ 809
42 4/5 gemeten

sql-data-warehouse-project

@DataWithBaraa

A comprehensive guide to building a modern data warehouse with SQL Server, including ETL processes, data modeling, and analytics.

MIT Commit 1 jaar geleden ★ 965
42 4/5 gemeten

dozer

@getdozer

Dozer is a real-time data movement tool that leverages CDC from various sources and moves data into various sinks.

AGPL-3.0 Commit 2 jaaren geleden ★ 1.578
42 4/5 gemeten

CQL

@CategoricalData

Categorical Query Language IDE

Commit 2 maanden geleden ★ 363
42 4/5 gemeten

orbital

@orbitalapi

Orbital automates integration between data sources (APIs, Databases, Queues and Functions). BFF's, API Composition and ETL pipelines that adapt as your specs change.

Commit 3 maanden geleden ★ 360
42 4/5 gemeten

aitrados-api

@aitrados

OHLC,news,economic event restfull api and WebSocket API for specifically designed for AI quantitative trading/training.Multiple Timeframes,Multiple-Symbols-Multiple-Timeframes

Apache-2.0 Commit 8 maanden geleden ★ 55
41 4/5 gemeten

monstache

@rwynn

a go daemon that syncs MongoDB to Elasticsearch in realtime. you know, for search.

MIT Commit 1 jaar geleden ★ 1.332
41 5/5 gemeten

efficient_data_processing_spark

@josephmachado

Code for "Efficient Data Processing in Spark" Course

Commit 3 maanden geleden ★ 393
41 4/5 gemeten

hellodata-be

@kanton-bern

The Open-Source Enterprise Data Platform in a single Portal

Commit 3 dagen geleden ★ 267
41 4/5 gemeten

ethereum-etl-airflow

@blockchain-etl

Airflow DAGs for exporting, loading, and parsing the Ethereum blockchain data. How to get any Ethereum smart contract into BigQuery https://towardsdatascience.com/how-to-get-any-ethereum-smart-contract-into-bigquery-in-8-mins-bab5db1fdeee

MIT Commit 1 jaar geleden ★ 440
40 4/5 gemeten

kiba

@thbar

Data processing & ETL framework for Ruby

Maintainer in de EU Commit 9 maanden geleden ★ 1.775
40 5/5 gemeten

diffsync

@networktocode

A utility library for comparing and synchronizing different datasets.

Commit 1 dag geleden ★ 189
40 5/5 gemeten

awesome-kafka

@infoslack

A list about Apache Kafka

Commit 5 maanden geleden ★ 593
39 5/5 gemeten

hudi-resources

@leesf

汇总Apache Hudi相关资料

Commit 6 maanden geleden ★ 555
39 4/5 gemeten

butterfree

@quintoandar

A tool for building feature stores.

Apache-2.0 Commit 8 maanden geleden ★ 319
39 4/5 gemeten

abc

@appbaseio

Power of appbase.io via CLI, with nifty imports from your favorite data sources

Apache-2.0 Commit 10 maanden geleden ★ 471
39 5/5 gemeten

goodreads_etl_pipeline

@san089

An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.

MIT Commit 7 jaaren geleden ★ 1.545
38 4/5 gemeten

memphis

@superstreamlabs

Memphis.dev is a highly scalable and effortless data streaming platform

Commit 7 maanden geleden ★ 3.445
38 4/5 gemeten

DataEngineeringProject

@damklis

Example end to end data engineering project.

Maintainer in de EU Commit 4 jaaren geleden ★ 1.434
38 4/5 gemeten

flupy

@olirice

Fluent data pipelines for python and your shell

Commit 2 maanden geleden ★ 196
38 4/5 gemeten

airbyte_serverless

@unytics

Airbyte made simple (no UI, no database, no cluster)

Maintainer in de EU Commit 1 jaar geleden ★ 196
37 4/5 gemeten

dataplane

@dataplane-app

Dataplane is an Airflow inspired unified data platform with additional data mesh and RPA capability to automate, schedule and design data pipelines and workflows. Dataplane is written in Golang with a React front end.

Commit 8 maanden geleden ★ 357
36 4/5 gemeten

pyper

@pyper-dev

Concurrent Python made simple

MIT Commit 2 jaaren geleden ★ 1.518
36 4/5 gemeten

omniparser

@jf-tech

omniparser: a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.

MIT Commit 2 jaaren geleden ★ 1.088
36 5/5 gemeten

etl2pcapng

@microsoft

Utility that converts an .etl file containing a Windows network packet capture into .pcapng format.

MIT Commit 1 jaar geleden ★ 740
36 4/5 gemeten

yobulkdev

@yobulkdev

🔥 🔥 🔥Open Source & AI driven Data Onboarding Platform:Free flatfile.com alternative

AGPL-3.0 Commit 3 jaaren geleden ★ 912
36 4/5 gemeten

mara-pipelines

@mara

A lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow

Maintainer in de EU Commit 3 jaaren geleden ★ 2.091
35 5/5 gemeten

harmony

@immunogenomics

Fast, sensitive and accurate integration of single-cell data with Harmony

Commit 4 maanden geleden ★ 672
35 4/5 gemeten

setl

@SETL-Framework

A simple Spark-powered ETL framework that just works 🍺

Apache-2.0 Commit 12 maanden geleden ★ 186
34 4/5 gemeten

getting-started

@singer-io

This repository is a getting started guide to Singer.

Commit 1 jaar geleden ★ 1.350
34 5/5 gemeten

klio

@spotify

Smarter data pipelines for audio.

Maintainer in de EU Commit 3 jaaren geleden ★ 876
34 5/5 gemeten

yuniql

@rdagumampan

Free and open source schema versioning and database migration made natively with .NET/6. NEW THIS MAY 2022! v1.3.15 released!

Maintainer in de EU Commit 2 jaaren geleden ★ 429
33 4/5 gemeten

bitcoin-etl

@blockchain-etl

ETL scripts for Bitcoin, Litecoin, Dash, Zcash, Doge, Bitcoin Cash. Available in Google BigQuery https://goo.gl/oY5BCQ

MIT Commit 1 jaar geleden ★ 461
32 5/5 gemeten

flock

@flock-lab

Flock: A Low-Cost Streaming Query Engine on FaaS Platforms

AGPL-3.0 Commit 3 jaaren geleden ★ 285
32 4/5 gemeten

spark-alchemy

@swoop-inc

Collection of open-source Spark tools & frameworks that have made the data engineering and data science teams at Swoop highly productive

Apache-2.0 Commit 12 maanden geleden ★ 192
32 5/5 gemeten

dud

@kevin-hanselman

A lightweight CLI tool for versioning data alongside source code and building data pipelines.

BSD-3-Clause Commit 1 jaar geleden ★ 220
31 5/5 gemeten

smooks

@smooks

An extensible Java framework for building event-driven applications that break up XML and non-XML data into chunks for data integration

Commit 10 maanden geleden ★ 421
31 5/5 gemeten

data-story

@ajthinking

A visual process builder

Maintainer in de EU Commit 1 jaar geleden ★ 217
31 5/5 gemeten

prefect-dataplatform

@anna-geller

Example repository showing how to build a data platform with Prefect, dbt and Snowflake

Maintainer in de EU Commit 4 jaaren geleden ★ 112
30 4/5 gemeten

active_workflow

@automaticmode

Polyglot workflows without leaving the comfort of your technology stack.

Zelf te hosten Commit 3 jaaren geleden ★ 864
30 4/5 gemeten

NeumAI

@NeumTry

Neum AI is a best-in-class framework to manage the creation and synchronization of vector embeddings at large scale.

Apache-2.0 Commit 3 jaaren geleden ★ 869
29 4/5 gemeten

SmartCode

@dotnetcore

SmartCode = IDataSource -> IBuildTask -> IOutput => Build Everything!!!

Apache-2.0 Commit 3 jaaren geleden ★ 577
29 4/5 gemeten