ARTICLE DETAIL

资讯详情

深耕郑州网站建设与运营推广的一线实战洞察。

Go 微服务监控告警:业务指标 vs 系统指标分层方案

Go 微服务监控告警:业务指标 vs 系统指标分层方案 Go 微服务监控告警业务指标 vs 系统指标分层方案监控不是 metrics 列表而是按层次梳理的策略。本文按 L7 应用 / L4 通道 / L3 基础设施分层讲透 Go 服务监控。一、三层监控原则业务指标L7订单数、登录数、失败率应用指标L6QPS、延迟、错误率系统指标L5CPU、内存、GC、goroutine二、业务指标案例订单服务varordersSuccessprometheus.NewCounterVec(prometheus.CounterOpts{Name:biz_orders_success_total},[]string{biz,channel},)varordersFailprometheus.NewCounterVec(prometheus.CounterOpts{Name:biz_orders_fail_total},[]string{biz,channel,reason},)业务侧failure reason枚举要有限。三、应用指标USE 原则Utilization资源占用Saturation饱和度Errors错误次数Go runtime 指标都有go_goroutines活跃 goroutinego_memstats_heap_inuse_bytes堆使用process_cpu_seconds_total进程占用四、RED 原则Rate每秒请求数Errors错误数Duration响应时长三者综合代表一个服务的健康度。五、Aggregation 关键技巧varhttpDurationprometheus.NewHistogramVec(prometheus.HistogramOpts{Name:http_duration_seconds,Buckets:[]float64{.005,.01,.025,.05,.1,.25,1,2.5,5,10},},[]string{method,route,status},)Buckets 决定 p99 计算。六、多 Label 引起的维度爆炸requestCount.WithLabelValues(GET,/a,200).Inc()requestCount.WithLabelValues(GET,/a,404).Inc()// 大量路由 大量 status 时TSDB 内存爆炸解决限制 status 到 4xx/5xx/2xx 大类路由拆分相似业务到 service_name label七、慢调用追踪histogram trace 联动start:time.Now()deferfunc(){httpDuration.WithLabelValues(...).Observe(time.Since(start).Seconds())}()1s 以上的慢查询记到 trace 系统进一步分析。八、灰度指标requestCount.WithLabelValues(method,route,status,version).Inc()通过 version label 区分 v1 v2 的性能。九、跨服务链路指标每个 service 都统计依赖的服务接口vardepLatencyprometheus.NewHistogramVec(prometheus.HistogramOpts{Name:dep_call_duration_seconds},[]string{target,method},)A→B 调用延迟。十、Alerting 策略-alert:HighErrorRateexpr:sum(rate(http_total{status~5..}[5m]))/sum(rate(http_total[5m]))0.05for:1m-alert:GoroutineSurgeexpr:go_goroutines10000针对每个 service 自定义规则。十一、生产实战指标分层层次采集方式频率存储业务显式埋点100%TSDB应用中间件100%TSDB系统runtime100%TSDB网络sidecar100%TSDBTraceOTel SDK5%trace store十二、踩坑清单Counter 重置32-bit 下 4.29B 上限 → 1.6y QPS 50% 写满 → 用 64-bitHistogram bucket 错过p95 跨 bucket 边界精度差export 失败处理promhttp export 失败不影响业务十三、总结与展望业务指标 RED USE 三大原则是监控核心。Go 端用client_golang体系完善配合 OTel 编织体系。未来AI Ops Observability 自动告警 异常检测 自动 RCA根因分析。十四、参考文献USE method (Brendan Gregg)Google SRE Bookprometheus 官方 docs
返回列表