splitting-datasets
|
它会碰到什么
扫了多少7 个文本文件,10 KB
它会碰到什么不碰外部(只输出文字)
命中总数0 处
命中统计严重 0 · 高 0 · 中 0 · 低 0
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Dataset Splitter
Positioning
Treat this skill as a narrow helper for partition strategy.
When to Use
Use this skill when:
- Prepare a dataset for machine learning model training.
- Create training, validation, and testing sets.
- Partition data to evaluate model performance.
Not For / Boundaries
- Full preprocessing-pipeline ownership: use
preprocessing-data-with-automated-pipelines - Leakage audits and prediction-time checks: use
ml-data-leakage-guard - Model training and tuning after the split: use
scikit-learn
Typical Outputs
- Partition strategy with ratios, random seeds, and stratification rules
- Notes on temporal or grouped split constraints
- Handoff guidance for leakage review and downstream training
Related Skills
preprocessing-data-with-automated-pipelinesfor the broader preprocessing sequenceml-data-leakage-guardto verify the split does not leak future or test information
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
它属于哪个仓库
星标★ 3,314
本站分层T1
该仓技能数258
原文件路径
bundled/skills/splitting-datasets/SKILL.md