跳到主要内容
知仓学习社ZHICANG

senior-data-engineer

Data engineering skill for building scalable data pipelines, ETL/ELT systems, and data infrastructure. Expertise in Python, SQL, Spark, Airflow, dbt…

读凭据读文件写文件执行命令联网严重 1 · 高危 0alirezarezvani/claude-skills

它会碰到什么

扫了多少9 个文本文件,275 KB
它会碰到什么读凭据读文件写文件执行命令联网
命中总数25 处
命中统计严重 1 · 高 0 · 中 23 · 低 0
逐条看命中(1 条严重或高危)
  • 严重 scripts/pipeline_orchestrator.py:118cred-paths
    's3': 'from airflow.providers.amazon.aws.operators.s3 import S3CreateBucketOperator',

这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。

技能内容

Senior Data Engineer

Production-grade data engineering skill for building scalable, reliable data systems.

Table of Contents

  1. [Trigger Phrases](#trigger-phrases)
  2. [Quick Start](#quick-start)
  3. [Workflows](#workflows)
  4. [Architecture Decision Framework](#architecture-decision-framework)
  5. [Tech Stack](#tech-stack)
  6. [Reference Documentation](#reference-documentation)
  7. [Troubleshooting](#troubleshooting)

Trigger Phrases

Activate this skill when you see:

Pipeline Design:

  • "Design a data pipeline for..."
  • "Build an ETL/ELT process..."
  • "How should I ingest data from..."
  • "Set up data extraction from..."

Architecture:

  • "Should I use batch or streaming?"
  • "Lambda vs Kappa architecture"
  • "How to handle late-arriving data"
  • "Design a data lakehouse"

Data Modeling:

  • "Create a dimensional model..."
  • "Star schema vs snowflake"
  • "Implement slowly changing dimensions"
  • "Design a data vault"

Data Quality:

  • "Add data validation to..."
  • "Set up data quality checks"
  • "Monitor data freshness"
  • "Implement data contracts"

Performance:

  • "Optimize this Spark job"
  • "Query is running slow"
  • "Reduce pipeline execution time"
  • "Tune Airflow DAG"

Quick Start

Core Tools

# Generate pipeline orchestration config
python scripts/pipeline_orchestrator.py generate \
  --type airflow \
  --source postgres \
  --destination snowflake \
  --schedule "0 5 * * *"

# Validate data quality
python scripts/data_quality_validator.py validate \
  --input data/sales.parquet \
  --schema schemas/sales.json \
  --checks freshness,completeness,uniqueness

# Optimize ETL performance
python scripts/etl_performance_optimizer.py analyze \
  --query queries/daily_aggregation.sql \
  --engine spark \
  --recommend

Workflows

→ See references/workflows.md for details

Architecture Decision Framework

Use this framework to choose the right approach for your data pipeline.

Batch vs Streaming

| Criteria | Batch | Streaming |

|----------|-------|-----------|

| Latency requirement | Hours to days | Seconds to minutes |

| Data volume | Large historical datasets | Continuous event streams |

| Processing complexity | Complex transformations, ML | Simple aggregations, filtering |

| Cost sensitivity | More cost-effective | Higher infrastructure cost |

| Error handling | Easier to reprocess | Requires careful design |

Decision Tree:

Is real-time insight required?
├── Yes → Use streaming
│   └── Is exactly-once semantics needed?
│       ├── Yes → Kafka + Flink/Spark Structured Streaming
│       └── No → Kafka + consumer groups
└── No → Use batch
    └── Is data volume > 1TB daily?
        ├── Yes → Spark/Databricks
        └── No → dbt + warehouse compute

Lambda vs Kappa Architecture

| Aspect | Lambda | Kappa |

|--------|--------|-------|

| Complexity | Two codebases (batch + stream) | Single codebase |

| Maintenance | Higher (sync batch/stream logic) | Lower |

| Reprocessing | Native batch layer | Replay from source |

| Use case | ML training + real-time serving | Pure event-driven |

When to choose Lambda:

  • Need to train ML models on historical data
  • Complex batch transformations not feasible in streaming
  • Existing batch infrastructure

When to choose Kappa:

  • Event-sourced architecture
  • All processing can be expressed as stream operations
  • Starting fresh without legacy systems

Data Warehouse vs Data Lakehouse

| Feature | Warehouse (Snowflake/BigQuery) | Lakehouse (Delta/Iceberg) |

|---------|-------------------------------|---------------------------|

| Best for | BI, SQL analytics | ML, unstructured data |

| Storage cost | Higher (proprietary format) | Lower (open formats) |

| Flexibility | Schema-on-write | Schema-on-read |

| Performance | Excellent for SQL | Good, improving |

| Ecosystem | Mature BI tools | Growing ML tooling |


Tech Stack

| Category | Technologies |

|----------|--------------|

| Languages | Python, SQL, Scala |

| Orchestration | Airflow, Prefect, Dagster |

| Transformation | dbt, Spark, Flink |

| Streaming | Kafka, Kinesis, Pub/Sub |

| Storage | S3, GCS, Delta Lake, Iceberg |

| Warehouses | Snowflake, BigQuery, Redshift, Databricks |

| Quality | Great Expectations, dbt tests, Monte Carlo |

| Monitoring | Prometheus, Grafana, Datadog |


Reference Documentation

1. Data Pipeline Architecture

See references/data_pipeline_architecture.md for:

  • Lambda vs Kappa architecture patterns
  • Batch processing with Spark and Airflow
  • Stream processing with Kafka and Flink
  • Exactly-once semantics implementation
  • Error handling and dead letter queues

2. Data Modeling Patterns

See references/data_modeling_patterns.md for:

  • Dimensional modeling (Star/Snowflake)
  • Slowly Changing Dimensions (SCD Types 1-6)
  • Data Vault modeling
  • dbt best practices
  • Partitioning and clustering

3. DataOps Best Practices

See references/dataops_best_practices.md for:

  • Data testing frameworks
  • Data contracts and schema validation
  • CI/CD for data pipelines
  • Observability and lineage
  • Incident response

Troubleshooting

→ See references/troubleshooting.md for details

想直接用这个技能?

本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。

它属于哪个仓库

星标★ 26,030
本站分层T1
该仓技能数846
原文件路径engineering-team/skills/senior-data-engineer/SKILL.md

同一个仓库里的其他技能

看这个仓库的全部 846 个技能

同名技能的其他版本

有 2 个不同仓库或目录里都有叫 senior-data-engineer 的技能。它们内容并不相同,别混用: