Data-integratie
89 projecten in Data & analyse
Kestra
@kestra-ioEvent-driven, language-agnostic platform to create, schedule, and monitor workflows. In code. Coordinate data pipelines and tasks such as ETL and ELT.
vector
@vectordotdevA high-performance observability data pipeline.
prefect
@PrefectHQPrefect is a workflow orchestration framework for building resilient data pipelines in Python.
whodb
@clideyWhere data access meets operational intelligence
steampipe
@turbotZero-ETL, infinite possibilities. Live query APIs, code & more with SQL. No DB required.
egeria
@odpiEgeria core
flink-cdc
@apacheFlink CDC is a streaming data integration tool
awesome-dbt
@HiflylabsA curated list of awesome dbt resources
airbyte
@airbytehqOpen-source data movement for ELT pipelines and AI agents — from APIs, databases & files to warehouses, lakes, and AI applications. Both self-hosted and Cloud.
hudi
@apacheUpserts, Deletes And Incremental Processing on Big Data.
reductstore
@reductstoreHigh-performance, time-indexed object storage for robotics and industrial IoT
jitsu
@jitsucomJitsu is an open-source Segment alternative. Fully-scriptable data ingestion engine for modern data teams. Set-up a real-time data pipeline in minutes, not days
od
@kokesČeská otevřená data
devlake
@apacheApache DevLake is an open-source dev data platform to ingest, analyze, and visualize the fragmented data from DevOps tools, extracting insights for engineering excellence, developer experience, and community growth.
proton
@timeplus-ioThe Fastest Unified Streaming SQL Engine in a Single C++ Binary. ⚡ Millisecond latency. 100+ GB/s throughput. Continuously compute real-time context from streams, logs, metrics, events, and CDC.
aws-sdk-pandas
@awspandas on AWS - Easy integration with Athena, Glue, Redshift, Timestream, Neptune, OpenSearch, QuickSight, Chime, CloudWatchLogs, DynamoDB, EMR, SecretManager, PostgreSQL, MySQL, SQLServer and S3 (Parquet, CSV, JSON and EXCEL).
meltano
@meltanoMeltano: the declarative code-first data integration engine that powers your wildest data and ML-powered product ideas. Say goodbye to writing, maintaining, and scaling your own API integrations.
desbordante-core
@DesbordanteDesbordante is a high-performance data profiler that is capable of discovering many different patterns in data using various algorithms. It also allows to run data cleaning scenarios using these algorithms. Desbordante has a console version and an easy-to-use web application.
duckle
@slothflowlabsOpen-source ETL/ELT you deploy on your own servers or cloud. Built on DuckDB: no-code/low-code visual pipelines or SQL, 385 components, dbt, CDC, data quality, reverse ETL, lineage, MCP for AI agents. No vendor cloud, no per-row billing.
squid-sdk
@subsquidTypeScript ETL toolkit for indexing Ethereum, Solana, and Substrate data, sourced from SQD Network.
Cookbook
@andkretThe Data Engineering Cookbook
hop
@apacheHop Orchestration Platform
gspread-pandas
@ParadigmllcRead and write Google Sheets as pandas DataFrames — column-matched appends, real dtypes, and layout detection for sheets that don't start at A1.
odd-platform
@opendatadiscoveryFirst open-source data discovery and observability platform. We make a life for data practitioners easy so you can focus on your business.
riko
@nerevuA Python stream processing engine modeled after Yahoo! Pipes
pyjanitor
@pyjanitor-devsClean APIs for data cleaning. Python implementation of R package Janitor
datacontract-cli
@datacontractEnforce Data Contracts
pathway
@pathwaycomPython ETL framework for stream processing, real-time analytics, LLM pipelines, and RAG.
lakeFS
@treeverselakeFS - Data version control for your data lake | Git for data
fluvio
@fluvio-community🦀 event stream processing for developers to collect and transform data in motion to power responsive data intensive applications.
chunjun
@DTStackA data integration framework
tis
@datavaneSupport agile Ontology DataOps Based on Flink, DataX and Flink-CDC with Web-UI
bacalhau
@bacalhau-projectCommunity-driven, simple, yet powerful framework for fast, cost-effective distributed Compute over Data.
soda-core
@sodadataData Contracts engine for the modern data stack. https://www.soda.io
doit
@pydoitCLI task management & automation tool
open-data-contract-standard
@bitol-ioHome of the Open Data Contract Standard (ODCS).
Knowledge-Graph-Tutorials-and-Papers
@heathersherryInsightful Tutorials and Papers about Knowledge Graphs
data-engineering-zoomcamp
@DataTalksClubData Engineering Zoomcamp is a free 9-week course on building production-ready data pipelines. Join the course here 👇🏼
zerocode
@authorjappszerocode-tdd is a community-developed, free, open-source, outcome-driven automated testing framework for Data Pipelines, ETL, REST API, Kafka(Data Streams), Databases and Load scenarios. all defined in simple JSON or YAML — with zero coding.
sling-cli
@slingdata-ioSling is a CLI tool that extracts data from a source storage/database and loads it in a target storage/database.
spark
@dataflintDrop-in replacement for Apache Spark UI
connect
@redpanda-dataFancy stream processing made operationally mundane
xonsh
@xonsh🐚 Python-powered shell. Full-featured, cross-platform and AI-friendly.
ingestr
@bruin-dataingestr is a CLI tool to copy data between any databases with a single command seamlessly.
wexflow
@aelassasWorkflow Automation Engine
awesome-etl
@pawlA curated list of awesome ETL frameworks, libraries, and software.
SDM-RDFizer
@SDM-TIBAn Efficient RML-Compliant Engine for Knowledge Graph Construction
cnpj-data-pipeline
@caiopizzolPipeline open-source que baixa e processa os dados da Receita Federal para PostgreSQL
quilt
@quiltdataQuilt is a Scientific Data Management Platform on AWS that helps teams and AI find, trust, and reuse data through deeply versioned, context-rich data packages.
awesome-node-based-uis
@xyflowA curated list with resources about node-based UIs
dataengineering-roadmap
@natayadevUn repositorio más con conceptos básicos, desafíos técnicos y recursos sobre ingeniería de datos en español 🧙✨
paperetl
@neuml📄 ⚙️ ETL processes for medical and scientific papers
public-datasets-pipelines
@GoogleCloudPlatformCloud-native, data onboarding architecture for Google Cloud Datasets
amphi-etl
@amphi-aivisual data prep powered by python
flow
@estuary🌊 Continuously synchronize the systems where your data lives, to the systems where you _want_ it to live, by managing your data flows with Estuary. 🌊
ChoETL
@CinchooETL framework for .NET (Parser / Writer for CSV, Flat, Xml, JSON, Key-Value, Parquet, Yaml, Avro formatted files)
dataforge
@Nuclear-MarmaladeFree open-source business data enrichment engine. The open-source alternative to Apollo, ZoomInfo, and Clearbit.
gusty
@pipeline-toolsMaking DAG construction easier
go-streams
@reugnA lightweight stream processing library for Go
pglogical
@2ndQuadrantLogical Replication extension for PostgreSQL 17, 16, 15, 14, 13, 12, 11, 10, 9.6, 9.5, 9.4 (Postgres), providing much faster replication than Slony, Bucardo or Londiste, as well as cross-version upgrades.
kuwala
@kuwala-ioKuwala is the no-code data platform for BI analysts and engineers enabling you to build powerful analytics workflows. We are set out to bring state-of-the-art data engineering tools you love, such as Airbyte, dbt, or Great Expectations together in one intuitive interface built with React Flow. In addition we provide third-party data into data science models and products with a focus on geospatial data. Currently, the following data connectors are available worldwide: a) High-resolution demographics data b) Point of Interests from Open Street Map c) Google Popular Times
koop
@koopjsTransform, query, and download geospatial data on the web.
qData
@qiantongtechqData is an open-source data governance and data development platform that integrates ETL, data development, metadata management, data quality, data assets, API services, and AI-powered data Q&A.
seatunnel-web
@apacheSeaTunnel is a distributed, high-performance data integration platform for the synchronization and transformation of massive data (offline & real-time).
complete-dbt-bootcamp-zero-to-hero
@zoltanctothSupplementary Materials for the The Complete dbt (Data Build Tool) Bootcamp Udemy course
awesome-data-catalogs
@opendatadiscovery📙 Awesome Data Catalogs and Observability Platforms.
practical-data-engineering
@ssp-dataPractical Data Engineering: A Hands-On Real-Estate Project Guide
goodreads_etl_pipeline
@san089An end-to-end GoodReads Data Pipeline for Building Data Lake, Data Warehouse and Analytics Platform.
kiba
@thbarData processing & ETL framework for Ruby
monstache
@rwynna go daemon that syncs MongoDB to Elasticsearch in realtime. you know, for search.
memphis
@superstreamlabsMemphis.dev is a highly scalable and effortless data streaming platform
DataEngineeringProject
@damklisExample end to end data engineering project.
mara-pipelines
@maraA lightweight opinionated ETL framework, halfway between plain scripts and Apache Airflow
omniparser
@jf-techomniparser: a native Golang ETL streaming parser and transform library for CSV, JSON, XML, EDI, text, etc.
flock
@flock-labFlock: A Low-Cost Streaming Query Engine on FaaS Platforms
pyper
@pyper-devConcurrent Python made simple
getting-started
@singer-ioThis repository is a getting started guide to Singer.
yobulkdev
@yobulkdev🔥 🔥 🔥Open Source & AI driven Data Onboarding Platform:Free flatfile.com alternative
klio
@spotifySmarter data pipelines for audio.
awesome-opensource-data-engineering
@gunnarmorlingAn Awesome List of Open-Source Data Engineering Projects
mycelial
@mycelialMove your data with ease.
data-engineer-roadmap
@datastacktvRoadmap to becoming a data engineer in 2021
active_workflow
@automaticmodePolyglot workflows without leaving the comfort of your technology stack.
sync-addons
@itpp-labs**Sync 🪬 Studio**
Data-Engineering-HowTo
@adilkhashA list of useful resources to learn Data Engineering from scratch
streamify
@ankurchavdaA data engineering project with Kafka, Spark Streaming, dbt, Docker, Airflow, Terraform, GCP and much more!
data-engineering-book
@oleg-agapovAccumulated knowledge and experience in the field of Data Engineering
open-data-etl-utility-kit
@ChicagoUse Pentaho's open source data integration tool (Kettle) to create Extract-Transform-Load (ETL) processes to update a Socrata open data portal. Documentation is available at http://open-data-etl-utility-kit.readthedocs.io/en/stable
RedditDataEngineering
@airscholarThis project provides a comprehensive data pipeline solution to extract, transform, and load (ETL) Reddit data into a Redshift data warehouse. The pipeline leverages a combination of tools and services including Apache Airflow, Celery, PostgreSQL, Amazon S3, AWS Glue, Amazon Athena, and Amazon Redshift.