nw-production-readiness
Monitoring, observability, operational procedures, CI/CD lessons learned, and quality gate definitions. Load when assessing production readiness or …
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
Production Readiness
Monitoring and Observability
Application Monitoring
- Performance: response time | throughput | latency percentiles (P50, P95, P99)
- Resources: CPU | memory | database connections | cache hit rates
- Errors: exception tracking | error rate trends | integration failure detection
- Business: KPI tracking | conversion funnels | feature usage | revenue impact
Infrastructure Monitoring
Server/container health and resource utilization | Network performance and connectivity | Storage capacity and I/O performance | Security event detection.
Alerting Tiers
| Tier | Condition | Response |
|------|-----------|----------|
| Page | Service down, data loss risk, security breach | Immediate response |
| Urgent | Error rate >2x baseline, latency SLA breach | Response within 15 min |
| Warning | Capacity >80%, error rate trending up | Response within 1 hour |
| Info | Deployment complete, metric threshold crossed | Review next business day |
Operational Procedures
Incident Response
- Detect: automated alerting identifies issue
- Triage: classify severity, assign responder
- Communicate: notify stakeholders per severity level
- Resolve: apply fix or rollback
- Review: post-incident review within 48 hours
- Improve: update runbooks and monitoring based on findings
Maintenance Procedures
Regular update and patching schedule | Backup verification (test restores quarterly) | Security vulnerability scanning (automated, weekly) | Performance baseline recalibration (after major changes).
Knowledge Transfer
Operational runbooks for common procedures | Architecture documentation with system diagrams | Deployment procedures and configuration management | Troubleshooting guides for known failure modes.
Quality Gates for Production Readiness
Before declaring production-ready, all must pass:
- [ ] All acceptance tests passing
- [ ] Unit coverage meets project standard (default: >= 80%)
- [ ] Integration tests validated
- [ ] Performance validated under realistic load
- [ ] Security scan completed (0 critical, 0 high)
- [ ] Monitoring and alerting configured
- [ ] Logging structured and searchable
- [ ] Rollback procedure documented and tested
- [ ] Runbook created for operational procedures
- [ ] On-call team trained on new feature
For CI/CD architecture lessons and measurement coupling pitfalls, see cicd-and-deployment skill.
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。
同名技能的其他版本
有 2 个不同仓库或目录里都有叫 nw-production-readiness 的技能。它们内容并不相同,别混用:
- nWave-ai/nWave — Monitoring, observability, operational procedures, CI/CD lessons learned, and quality gate