Instellingen

Plug-ins

Data Engineering Copilot

Debug production data systems

Plug-in installeren

Data Engineering Copilot helps you design, debug, recover, and improve production data systems across Spark, Databricks, Delta Lake, SQL, CDC, Structured Streaming, Airflow, backfills, data quality, observability, and lakehouse architectures. It can diagnose failed or slow pipelines, reason about duplicate and late data, plan safe replay and backfill strategies, review checkpoint and streaming-state risks, tune Spark from execution evidence, and create production-oriented SQL, PySpark, and orchestration patterns. It prioritizes correctness, recoverability, SLA, security, and cost before scaling or redesigning.

Skill

Informatie

Functies
Design reliable batch, CDC, streaming, and lakehouse architectures, Debug failed, slow, duplicated, stale, or backlogged production pipelines, Plan safe replay, backfill, recovery, and checkpoint strategies, Review Delta MERGE logic for keys, ordering, duplicates, updates, and deletes, Diagnose Spark and Databricks performance using plans, shuffle, spill, skew, and file evidence, Design Structured Streaming state, watermark, checkpoint, and sink behavior, Create production-oriented SQL, PySpark, Databricks, and Airflow patterns, Build data-quality gates, reconciliation, freshness checks, and operational runbooks, Evaluate retry safety, idempotency, schema evolution, and downstream contracts, Use current official docs for Spark, Databricks, Delta Lake, Airflow, and runtime-specific behavior
Ontwikkelaar
Krishna Sathvik
Categorie
Data & Analytics
Versie
0.1.0