preprocessing-data-with-automated-pipelines
|
它会碰到什么
扫了多少8 个文本文件,45 KB
它会碰到什么执行命令写文件读文件
命中总数15 处
命中统计严重 0 · 高 1 · 中 14 · 低 0
逐条看命中(1 条严重或高危)
- 高
scripts/pipeline.py:78exec-spawnresult = subprocess.run(
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Data Preprocessing Pipeline
Positioning
Use this skill as the direct owner for ML input-preparation pipelines.
It covers preprocessing-heavy tasks where the requested deliverable is a repeatable pipeline for cleaning, encoding, transforming, and validating input data.
When to Use
Use this skill when:
- Prepare raw data for machine learning models.
- Automate data cleaning and transformation processes.
- Implement a robust ETL (Extract, Transform, Load) pipeline.
Not For / Boundaries
- Whole-task ML ownership: use
scikit-learnorml-pipeline-workflow - Leakage and prediction-time auditing: use
ml-data-leakage-guard - Grouped scientific preprocessing with stronger methodological constraints: use
scientific-data-preprocessing
Typical Outputs
- A preprocessing pipeline plan or implementation sketch
- Clear sequencing for clean, encode, transform, and validate steps
- Notes that identify where leakage review, training, or evaluation should be run next
Related Skills
ml-data-leakage-guardbefore trusting fitted preprocessing stepssplitting-datasetswhen the next narrow problem is partition strategy
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
星标★ 3,314
本站分层T1
该仓技能数258
原文件路径
bundled/skills/preprocessing-data-with-automated-pipelines/SKILL.md