gke-observability
>-
它会碰到什么
这一栏是扫描器报的事实,不是结论。命中多不等于有毒(安全工具、规则库、示例脚本本来就会包含危险写法),命中少也不等于干净。它和你手上的凭据、文件、网络有什么关系,需要你自己看。
技能内容
GKE Observability
This reference covers monitoring, logging, and metrics configuration for GKE.
The golden path enables comprehensive observability including control-plane
metrics.
> MCP Tools: get_cluster, list_k8s_events, get_k8s_logs,
> get_k8s_cluster_info, describe_k8s_resource. CLI-only: `gcloud
> container clusters update --monitoring=..., gcloud logging read`
Golden Path Observability Defaults
Setting | Golden Path Value | Notes
--------------------------------------------------- | --------------------------------------------------------------------------------------------------------------------------------------------------- | -----
loggingConfig components | SYSTEM_COMPONENTS, WORKLOADS | Full workload logging
monitoringConfig components | SYSTEM_COMPONENTS, STORAGE, POD, DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, JOBSET, CADVISOR, KUBELET, DCGM, APISERVER, SCHEDULER, CONTROLLER_MANAGER | Full suite including control-plane
managedPrometheusConfig.enabled | true | Google-managed Prometheus
advancedDatapathObservabilityConfig.enableMetrics | true | Dataplane V2 flow metrics
loggingService | logging.googleapis.com/kubernetes | Cloud Logging
monitoringService | monitoring.googleapis.com/kubernetes | Cloud Monitoring
Control-Plane Metrics (Golden Path Addition)
The golden path adds three control-plane monitoring components not present in
default clusters:
| Component | What It Monitors |
| -------------------- | ---------------------------------------------------------------------- |
| APISERVER | API server request latency, error rates, admission webhook performance |
| SCHEDULER | Scheduling latency, pending pods, scheduling failures |
| CONTROLLER_MANAGER | Controller work queue depth, reconciliation latency |
These are critical for diagnosing cluster-level issues (slow API responses,
scheduling delays, stuck controllers).
Enabling Full Monitoring
Say this whenever you hand over a --monitoring command:
- Control-plane metrics are NOT enabled by default. State this outright in
your answer — do not leave it implied by the fact that you are supplying an
enable command. API_SERVER, SCHEDULER, and CONTROLLER_MANAGER are off
on every new cluster and collect nothing until explicitly turned on, and the
same is true of DCGM, CADVISOR, KUBELET, and kube-state (POD,
DEPLOYMENT, STATEFULSET, DAEMONSET, HPA, STORAGE, JOBSET).
SYSTEM is the only package on by default. A user asking "why are there no
API server metrics" has almost always simply never enabled them.
- The flag replaces, it does not append. The set supplied to
--monitoring
overrides the previous setting entirely, so omitting a component silently
turns it off. Always pass the full desired list, and always include SYSTEM
— it cannot be disabled while monitoring is on, and never on Autopilot.
- These metrics bill per sample ingested via Managed Service for
Prometheus. Enabling the full suite on a large cluster is a real cost
increase; mention it rather than presenting the list as free.
> **The gcloud flag and the API field use different spellings for the same
> components.** Do not copy names between them:
>
> Component | gcloud --monitoring= | monitoringConfig API enum
> ---------------- | ---------------------- | ---------------------------
> System | SYSTEM | SYSTEM_COMPONENTS
> API server | API_SERVER | APISERVER
> Controller mgr | CONTROLLER_MANAGER | CONTROLLER_MANAGER
>
> The remaining components share a spelling. Using an API enum in the CLI flag
> (or the reverse) fails the command — this is a common and confusing error.
# Enable golden path monitoring suite
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM,API_SERVER,SCHEDULER,CONTROLLER_MANAGER,STORAGE,POD,DEPLOYMENT,STATEFULSET,DAEMONSET,HPA,JOBSET,CADVISOR,KUBELET,DCGM \
--quiet
# Enable Managed Prometheus
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-managed-prometheus \
--quiet
# Enable Dataplane V2 observability metrics
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--enable-dataplane-v2-flow-observability \
--quiet
Managed Prometheus
Golden path enables Google Managed Prometheus for metrics collection and
querying.
Querying metrics:
- Use Cloud Monitoring Metrics Explorer in the console
- Use PromQL via the Prometheus UI or API
- Grafana dashboards via Managed Grafana
Key GKE metrics:
| Metric | Source | Use |
| -------------------------------------------------- | ------------------ | ---------------------- |
| container_cpu_usage_seconds_total | cAdvisor | Pod CPU usage |
| container_memory_working_set_bytes | cAdvisor | Pod memory usage |
| kube_pod_status_phase | kube-state-metrics | Pod lifecycle |
| apiserver_request_duration_seconds | API Server | Control plane latency |
| scheduler_scheduling_attempt_duration_seconds | Scheduler | Scheduling performance |
| kubernetes.io/node/cpu/core_usage_time | Cloud Monitoring | Node CPU |
| DCGM_FI_DEV_GPU_UTIL | DCGM | GPU utilization |
Live Resource Usage (kubectl-only)
No MCP or gcloud equivalent exists for live resource usage. Use kubectl top:
kubectl top pods --all-namespaces --sort-by=cpu
kubectl top nodes
kubectl top pods --containers -n <NAMESPACE> # per-container breakdown
Cloud Logging (gcloud-only)
Querying cluster logs (no MCP equivalent — use gcloud logging read):
# System component logs
gcloud logging read \
'resource.type="k8s_cluster" AND resource.labels.cluster_name="<CLUSTER_NAME>"' \
--project <PROJECT_ID> --limit 50 \
--quiet
# Workload logs for a specific namespace
gcloud logging read \
'resource.type="k8s_container" AND resource.labels.cluster_name="<CLUSTER_NAME>" AND resource.labels.namespace_name="<NAMESPACE>"' \
--project <PROJECT_ID> --limit 50 \
--quiet
# Audit logs (who did what)
gcloud logging read \
'resource.type="k8s_cluster" AND logName:"cloudaudit.googleapis.com"' \
--project <PROJECT_ID> --limit 50 \
--quiet
Diagnostic Settings
For security monitoring and troubleshooting, enable control-plane audit logs:
# View current logging config
gcloud container clusters describe <CLUSTER_NAME> --region <REGION> \
--format="yaml(loggingConfig)" \
--quiet
Alerting
Set up alerts for critical conditions:
Condition | Metric | Threshold
----------------------- | --------------------------------------------------- | ---------
High API server latency | apiserver_request_duration_seconds | P99 > 5s
Pod crash loops | kube_pod_container_status_restarts_total | > 5 in 10min
Node not ready | kube_node_status_condition | condition=Ready, status!=True
High GPU utilization | DCGM_FI_DEV_GPU_UTIL | > 95% sustained
PVC near capacity | kubelet_volume_stats_used_bytes / capacity | > 85%
Scheduling failures | scheduler_schedule_attempts_total{result="error"} | > 0
> Prerequisite: The kube_* series above (e.g., kube_pod_status_phase,
> kube_pod_container_status_restarts_total, kube_node_status_condition)
> come from kube-state-metrics, which GKE does not collect by default.
> Deploy the Managed Prometheus kube-state-metrics package first.
Proposing Dashboards & Alerts (Production Rules)
When designing or proposing alerting and dashboard strategies for GKE:
- Always explicitly name Google Cloud Monitoring as the platform to
implement these alerts and dashboards.
- Always include API server latency (via
apiserver_request_duration_seconds metric) on the dashboard as a critical
indicator of control plane health, alongside node CPU/Memory and pod crash
loops.
Node Health (Production Rules)
A comprehensive assessment of node health relies on analyzing these two metrics together:
kubernetes.io/node/status_condition(filtered bystatus_condition="Ready"): Use this to track healthy nodes. Note that it will only report values for nodes that have successfully bootstrapped.compute.googleapis.com/instance_group/size(filtered byinstance_group_name="gke-<cluster_name>-.*"): Use this to track the total number of nodes in a specific cluster. Note that it does not differentiate between healthy and unhealthy nodes.
Cost Considerations
Monitoring and logging have associated costs:
- Cloud Logging: Charged per GiB ingested beyond free tier (50
GiB/project/month)
- Cloud Monitoring: Free for GKE system metrics; custom metrics charged
per time series
- Managed Prometheus: Charged per samples ingested
To reduce costs in non-production:
# Reduce to system-only monitoring
gcloud container clusters update <CLUSTER_NAME> --region <REGION> \
--monitoring=SYSTEM \
--quiet
Distributed Tracing & Continuous Profiling (Recommended)
Not golden path defaults — recommended for production microservice
architectures and performance-sensitive workloads.
- Cloud Trace: Add OpenTelemetry SDK to your app with the
opentelemetry-operations-go (or equivalent) exporter. Traces appear in
Cloud Trace console. Identifies cross-service latency bottlenecks.
- Cloud Profiler: Add the Cloud Profiler agent to your app. Profiles CPU
and memory usage in production with low overhead. Identifies hotspots and
compares across versions.
Recent additions:
- Managed OpenTelemetry for GKE (Preview): Managed in-cluster OTLP
endpoint plus auto-instrumentation for traces, metrics, and logs. Requires
GKE 1.34.1-gke.2178000+; enable with `gcloud beta container clusters
update ... --managed-otel-scope=COLLECTION_AND_INSTRUMENTATION_COMPONENTS`.
- PSI (Pressure Stall Information) metrics: cAdvisor
container_pressure_{cpu,memory,io}_{waiting,stalled}_seconds_total series
(beta in Kubernetes 1.34) can be collected via a Managed Prometheus
ClusterNodeMonitoring resource; GKE's documented collection path requires
GKE 1.35+.
LQL Query Examples
Common Logging Query Language patterns for GKE troubleshooting:
# Error logs for a specific container
resource.type="k8s_container" AND resource.labels.container_name="my-app" AND severity>=ERROR
# OOMKilled events
resource.type="k8s_event" AND jsonPayload.reason="OOMKilling"
# Pod scheduling failures
resource.type="k8s_event" AND jsonPayload.reason="FailedScheduling"
# Audit logs (who did what)
resource.type="k8s_cluster" AND logName:"cloudaudit.googleapis.com"
Supporting Links
想直接用这个技能?
本站把开放许可(MIT / Apache 等)的技能按仓库打包整理到网盘,点一下转存到你自己的网盘,不用一个个从 GitHub 拉。许可未声明的技能只给原始仓库链接,不打包。