mirror of
https://github.com/milvus-io/milvus.git
synced 2026-07-21 10:15:43 +00:00
issue: #51493 ### What #50221 split the `status` label values of `milvus_proxy_req_count` / `milvus_proxy_grpc_latency` in place. Released in **v2.6.19**, that silently zeroes every existing dashboard panel and alert rule matching `status="fail"` / `status="rejected"` — an alert that stops firing rather than erroring. This keeps the classification but moves it onto its own dimension, restoring the status domain to what <= v2.6.18 emitted: ``` status: success | fail | rejected | retry | total | abandon (as before v2.6.19) cause: user | system | cancel | na (new) ``` | v2.6.19 / v2.6.20 | this PR | |---|---| | `status="fail_input"` | `status="fail", cause="user"` | | `status="fail_system"` | `status="fail", cause="system"` | | `status="rejected_user"` | `status="rejected", cause="user"` | | `status="rejected_system"` | `status="rejected", cause="system"` | | `status="cancel"` | `status="fail"` or `"rejected"`, `cause="cancel"` | | `success` / `retry` / `total` / `abandon` | unchanged, `cause="na"` | ### Why a label instead of new status values - **Old queries work unchanged and count each request once.** `status="fail"` is again every hard failure; Prometheus aggregates over `cause` for free. That rollup is what a label dimension *is* — the split forces every consumer who just wants "how many failures" to hand-write `status=~"fail_.*"`. - **Cardinality is unchanged.** `cause` is functionally dependent on the outcome, so the realized `(status, cause)` pairs are exactly the series the split already produces. Additive, not multiplicative. - **New consumers lose nothing**: alert on `status="fail", cause="system"`; route `cause="user"` to the owning application team. The alternative — emitting old and new `status` values side by side during a transition — was rejected: it defers the break rather than removing it (the old values must still be dropped eventually, forcing the same migration later), and meanwhile double-counts every request in unfiltered aggregations. `cancel` is folded into `cause` for the same reason: before v2.6.19 client cancellations were counted as `fail` (cancel in the response status) or `rejected` (cancel at the interceptor), so promoting it to a `status` value quietly makes `status="fail"` under-count versus older releases. As a cause it stays excludable via `cause!="cancel"` without redefining `status`. ### Also in this PR - The 7 `snapshot_impl` sites that emitted a bare `fail` with no classification now report their real cause via `failMetricLabel(err)` (e.g. `snapshot_metadata_uri is required` is `cause="user"`, not a system fault). Their `status` is unchanged. - `deployments/monitor/grafana/milvus-dashboard.json`: the "Faild Request Rate" panel goes back to `status="fail"` — a verbatim revert of what #50221 changed, which is the compatibility claim demonstrated. - Dev docs and stale comments that named the old label values. ### Verification - `TestParseMetricLabelStatusDomainIsStable` pins the status domain, so re-splitting it fails CI rather than shipping. - Adding a label makes every `WithLabelValues` site arity-sensitive **at runtime, not compile time** — a missed site panics in production. Unit tests do not reach all of them (coverage shows `GetRestoreSnapshotState`, `ListRestoreSnapshotJobs`, `PinSnapshotData`, `UnpinSnapshotData` at 0%), so all **58 emit sites were verified statically via AST**, not only the ones tests happen to hit. - Passing: `pkg/metrics`, `pkg/util/requestutil`, `pkg/util/merr`, `internal/distributed/proxy` (incl. `httpserver`, which covers the REST emit path), and the `internal/proxy` snapshot / metric-label tests. `golangci-lint` clean on every touched package. - Pre-existing and unrelated: `service_test.go` bind failures (`listen tcp :19529: address already in use`) reproduce identically on unmodified master — a local process holds the port. ### Rollout Should be cherry-picked to 2.6 so the window in which the fine-grained `status` values exist stays confined to v2.6.19–v2.6.20. Users who already adopted the new values on those two releases need a one-time query change (`status="fail_system"` -> `status="fail", cause="system"`); that is a known, bounded set, and preferable to leaving every pre-2.6.19 dashboard silently broken. 🤖 Generated with [Claude Code](https://claude.com/claude-code) Signed-off-by: zhenshan.cao <zhenshan.cao@zilliz.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
240 lines
8.3 KiB
Go
240 lines
8.3 KiB
Go
// Licensed to the LF AI & Data foundation under one
|
|
// or more contributor license agreements. See the NOTICE file
|
|
// distributed with this work for additional information
|
|
// regarding copyright ownership. The ASF licenses this file
|
|
// to you under the Apache License, Version 2.0 (the
|
|
// "License"); you may not use this file except in compliance
|
|
// with the License. You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
package metrics
|
|
|
|
import (
|
|
// #nosec
|
|
_ "net/http/pprof"
|
|
|
|
"github.com/prometheus/client_golang/prometheus"
|
|
)
|
|
|
|
const (
|
|
milvusNamespace = "milvus"
|
|
|
|
AbandonLabel = "abandon"
|
|
SuccessLabel = "success"
|
|
FailLabel = "fail"
|
|
CancelLabel = "cancel"
|
|
TotalLabel = "total"
|
|
RetryLabel = "retry"
|
|
RejectedLabel = "rejected"
|
|
|
|
// Values of the "cause" label, a dimension orthogonal to "status" that names
|
|
// the responsible party for a failed request, so monitoring can tell a
|
|
// user-input error (the caller must fix the request) apart from an internal
|
|
// system error (operators must intervene). It is a separate label rather
|
|
// than extra "status" values on purpose: the coarse status stays a valid
|
|
// query on its own ("fail" is still every hard failure, aggregated over
|
|
// cause by Prometheus), so dashboards written against it keep working.
|
|
// Cause is functionally dependent on the outcome, so the realized
|
|
// (status, cause) pairs are additive, not the product of both domains.
|
|
CauseUser = "user" // the request itself is at fault: bad arguments, missing auth, no such collection
|
|
CauseSystem = "system" // Milvus is at fault: component failure, IO error, internal bug
|
|
CauseCancel = "cancel" // neither party: the client gave up before the request completed
|
|
CauseNA = "na" // no cause applies: the request did not hard-fail
|
|
|
|
HybridSearchLabel = "hybrid_search"
|
|
|
|
InsertLabel = "insert"
|
|
DeleteLabel = "delete"
|
|
UpsertLabel = "upsert"
|
|
SearchLabel = "search"
|
|
QueryLabel = "query"
|
|
UpsertQueryLabel = "upsert_query"
|
|
DeleteQueryLabel = "delete_query"
|
|
ReQueryLabel = "requery"
|
|
CacheHitLabel = "hit"
|
|
CacheMissLabel = "miss"
|
|
TimetickLabel = "timetick"
|
|
AllLabel = "all"
|
|
)
|
|
|
|
const (
|
|
PreferredNodeHitLabel = "hit"
|
|
PreferredNodeMissLabel = "miss"
|
|
PreferredNodeUnavailableLabel = "unavailable"
|
|
PreferredNodeRejectedLabel = "rejected"
|
|
)
|
|
|
|
const (
|
|
UnissuedIndexTaskLabel = "unissued"
|
|
InProgressIndexTaskLabel = "in-progress"
|
|
FinishedIndexTaskLabel = "finished"
|
|
FailedIndexTaskLabel = "failed"
|
|
RecycledIndexTaskLabel = "recycled"
|
|
|
|
// Note: below must matchcommonpb.SegmentState_name fields.
|
|
SealedSegmentLabel = "Sealed"
|
|
GrowingSegmentLabel = "Growing"
|
|
FlushedSegmentLabel = "Flushed"
|
|
FlushingSegmentLabel = "Flushing"
|
|
DroppedSegmentLabel = "Dropped"
|
|
|
|
StreamingDataSourceLabel = "streaming"
|
|
BulkinsertDataSourceLabel = "bulkinsert"
|
|
CompactionDataSourceLabel = "compaction"
|
|
|
|
Leader = "OnLeader"
|
|
FromLeader = "FromLeader"
|
|
|
|
HookBefore = "before"
|
|
HookAfter = "after"
|
|
HookMock = "mock"
|
|
|
|
ReduceSegments = "segments"
|
|
ReduceShards = "shards"
|
|
|
|
BatchReduce = "batch_reduce"
|
|
|
|
FunctionChainLevelL0 = "l0"
|
|
|
|
Pending = "pending"
|
|
Executing = "executing"
|
|
Done = "done"
|
|
|
|
ImportStagePending = "pending"
|
|
ImportStagePreImport = "preimport"
|
|
ImportStageImport = "import"
|
|
ImportStageStats = "stats"
|
|
ImportStageBuildIndex = "build_index"
|
|
ImportStageWaitL0Import = "wait_l0_import"
|
|
|
|
compactionTypeLabelName = "compaction_type"
|
|
isVectorFieldLabelName = "is_vector_field"
|
|
segmentPruneLabelName = "segment_prune_label"
|
|
stageLabelName = "compaction_stage"
|
|
nodeIDLabelName = "node_id"
|
|
nodeHostLabelName = "node_host"
|
|
statusLabelName = "status"
|
|
causeLabelName = "cause"
|
|
indexTaskStatusLabelName = "index_task_status"
|
|
msgTypeLabelName = "msg_type"
|
|
collectionIDLabelName = "collection_id"
|
|
fieldIDLabelName = "field_id"
|
|
channelNameLabelName = "channel_name"
|
|
functionLabelName = "function_name"
|
|
queryTypeLabelName = "query_type"
|
|
chainLevelLabelName = "chain_level"
|
|
collectionName = "collection_name"
|
|
databaseLabelName = "db_name"
|
|
ResourceGroupLabelName = "rg"
|
|
indexName = "index_name"
|
|
isVectorIndex = "is_vector_index"
|
|
segmentStateLabelName = "segment_state"
|
|
segmentLevelLabelName = "segment_level"
|
|
segmentIsSortedLabelName = "segment_is_sorted"
|
|
segmentStorageVersionLabelName = "segment_storage_version"
|
|
segmentFormatLabelName = "segment_format"
|
|
usernameLabelName = "username"
|
|
roleNameLabelName = "role_name"
|
|
cacheNameLabelName = "cache_name"
|
|
cacheStateLabelName = "cache_state"
|
|
dataSourceLabelName = "data_source"
|
|
dataTypeLabelName = "data_type"
|
|
importStageLabelName = "import_stage"
|
|
requestScope = "scope"
|
|
fullMethodLabelName = "full_method"
|
|
reduceLevelName = "reduce_level"
|
|
reduceType = "reduce_type"
|
|
lockName = "lock_name"
|
|
lockSource = "lock_source"
|
|
lockType = "lock_type"
|
|
lockOp = "lock_op"
|
|
loadTypeName = "load_type"
|
|
pathLabelName = "path"
|
|
cgoNameLabelName = `cgo_name`
|
|
cgoTypeLabelName = `cgo_type`
|
|
queueTypeLabelName = `queue_type`
|
|
poolNameLabelName = "pool_name"
|
|
outcomeLabelName = "outcome"
|
|
|
|
// model function/UDF labels
|
|
functionTypeName = "function_type_name"
|
|
functionProvider = "function_provider"
|
|
functionName = "function_name"
|
|
|
|
// entities label
|
|
LoadedLabel = "loaded"
|
|
NumEntitiesAllLabel = "all"
|
|
|
|
TaskTypeLabel = "task_type"
|
|
TaskStateLabel = "task_state"
|
|
|
|
filesystemKeyLabelName = "fs"
|
|
reasonLabelName = "reason"
|
|
)
|
|
|
|
var (
|
|
// buckets involves durations in milliseconds,
|
|
// [1 2 4 8 16 32 64 128 256 512 1024 2048 4096 8192 16384 32768 65536 1.31072e+05]
|
|
buckets = prometheus.ExponentialBuckets(1, 2, 18)
|
|
|
|
// subMsBuckets extends buckets with sub-millisecond boundaries for histograms
|
|
// whose observations can fall below 1ms (e.g. in-memory or queue-wait paths).
|
|
// [0.1 0.25 0.5 1 2 4 8 16 32 64 128 256 512 1024 2048 4096 8192 16384 32768 65536 1.31072e+05]
|
|
subMsBuckets = append([]float64{0.1, 0.25, 0.5}, prometheus.ExponentialBuckets(1, 2, 18)...)
|
|
|
|
// longTaskBuckets provides long task duration in milliseconds
|
|
longTaskBuckets = []float64{1, 100, 500, 1000, 5000, 10000, 20000, 50000, 100000, 250000, 500000, 1000000, 3600000, 5000000, 10000000} // unit milliseconds
|
|
|
|
// size provides size in byte
|
|
sizeBuckets = []float64{10000, 100000, 1000000, 100000000, 500000000, 1024000000, 2048000000, 4096000000, 10000000000, 50000000000} // unit byte
|
|
|
|
NumNodes = prometheus.NewGaugeVec(
|
|
prometheus.GaugeOpts{
|
|
Namespace: milvusNamespace,
|
|
Name: "num_node",
|
|
Help: "number of nodes and coordinates",
|
|
}, []string{nodeIDLabelName, roleNameLabelName})
|
|
|
|
LockCosts = prometheus.NewGaugeVec(
|
|
prometheus.GaugeOpts{
|
|
Namespace: milvusNamespace,
|
|
Name: "lock_time_cost",
|
|
Help: "time cost for various kinds of locks",
|
|
}, []string{
|
|
lockName,
|
|
lockSource,
|
|
lockType,
|
|
lockOp,
|
|
})
|
|
|
|
metricRegisterer prometheus.Registerer
|
|
)
|
|
|
|
// GetRegisterer returns the global prometheus registerer
|
|
// metricsRegistry must be call after Register is called or no Register is called.
|
|
func GetRegisterer() prometheus.Registerer {
|
|
if metricRegisterer == nil {
|
|
return prometheus.DefaultRegisterer
|
|
}
|
|
return metricRegisterer
|
|
}
|
|
|
|
// Register serves prometheus http service
|
|
// Should be called by init function.
|
|
func Register(r prometheus.Registerer) {
|
|
r.MustRegister(NumNodes)
|
|
r.MustRegister(LockCosts)
|
|
r.MustRegister(BuildInfo)
|
|
r.MustRegister(RuntimeInfo)
|
|
r.MustRegister(ThreadNum)
|
|
r.MustRegister(ThreadCPUActiveNumByPool)
|
|
metricRegisterer = r
|
|
}
|