Files
milvus/internal/datacoord/import_meta.go
T
e2787d3981 enhance: standardize error handling on merr + Sys/Input classification (#50221)
issue: #47420

## What this PR does

Project-wide migration of raw `fmt.Errorf` / `errors.New` in function
bodies onto
the `merr` framework, plus the Sys-vs-Input error classification and the
machinery it drives (retriability, fine-grained metrics, segcore
unification),
plus the convention docs and a linter that keeps it from regressing.

Scope: storage, proxy, coordinators (root/data/query), query node, data
node,
`pkg/util` & `internal/util`, expression parser, message queue,
streaming, and
misc packages. Bare raw-error usages went from ~3000 to a ~340 allowlist
(package-level sentinels / build-tag / test sites).

---

## How to review this PR

It is large but the vast majority is mechanical. Changes fall into three
tiers;
spend review budget on Part 2 and Part 3.

### Part 1 — Mechanical standardization (low risk, verify by rule)

Each converted call follows one of a small fixed set of rules. To
review, check
that each site obeys the matching rule rather than reading every line:

| Pattern | Rule |
|---|---|
| `fmt.Errorf("...")` originating a new error | →
`merr.WrapErrXxxMsg("...")` with a code matching the failure's meaning |
| Adding context to an existing typed error | → `merr.Wrap(err, "...")`
/ `merr.Wrapf(...)` — **preserves** the inner code (never `WrapErr*Err`,
which overwrites it) |
| Errors inside the streaming subsystem | → `status.New*` factories
(StreamingError), **not** merr — this is the component-internal dialect
(see `docs/dev/error_handling_guide.md`) |
| Low-level / control-flow signal caught by `errors.Is` | → kept as a
package-level `errors.New` sentinel (lowercase, same-package) |

Conventions are documented in `docs/dev/error_handling_guide.md`
(how-to) and
`docs/dev/error_sentinel_convention.md` (rules + audit). A
`gocritic`/`ruleguard`
rule (`rawmerrerror`, in `rules.go`) enforces "no raw `return
errors.New/fmt.Errorf`"
under `make verifiers`.

### Part 2 — Behavior changes (review these closely)

These are the sites where the wire contract or runtime behavior changes,
not just
the source text. Listed by category; representative locations given,
full set in
the diff.

**A. gRPC wire-code shifts: `UnexpectedError(1)/Code 65535` → typed
code.**
Where a handler previously returned a raw error (collapsed to
`Code=65535` on the
wire), it now returns a typed merr, so the client sees a real code. The
most
common shift is to `IllegalArgument(5)/Code 1100` (ParameterInvalid).
Touch
points include datanode task handlers (CreateTask/Query/Drop), proxy
Upsert,
querynode GetMetrics, datacoord CreateIndex, httpserver query-response
builder,
and typeutil schema validation. One code refinement: an index-param
validation
moved `1100` → `1101` (ParameterMissing). **Client/SDK assertions and
any code
that switched on `Code=65535` for these paths must be re-checked** (the
go_client
e2e assertions were already aligned in this PR).

**B. Prometheus `status` label contract change (externally visible).**
The proxy metric's coarse `fail` / `rejected` values are split into
`fail_input` / `fail_system` and `rejected_user` / `rejected_system` (in
`requestutil.ParseMetricLabel`; auth/privilege rejections count as
`rejected_user`), so dashboards can attribute a failure to caller vs
operator.
**Dashboards/alerts querying `status="fail"` must migrate to
`status=~"fail_.*"`, and `status="rejected"` to
`status=~"rejected_.*"`.** The
in-repo Grafana dashboard is already migrated; external dashboards built
on the
old values silently go empty after upgrade. This is the one change that
requires an ops-side migration.

**C. Retriability semantics.**
- C1: `merr.Status(err)` now forces `Retriable=false` when the error is
an
`InputError` — a malformed request can never succeed on blind retry, so
clients
never get the self-contradictory "your input is wrong but you may
retry".
- C2: `retry.Do` short-circuits an `InputError` (non-retriable) — **but
only when
  the caller did not pass a `RetryErr` predicate**. The check is an
`if c.isRetryErr != nil { ... } else if InputError { ... }` *mutually
exclusive*
branch (`pkg/util/retry/retry.go`): an explicit `RetryErr` takes
precedence and
bypasses the InputError abort. `retry.Handle` deliberately does **not**
apply
the InputError abort (its callers signal abort via `shouldRetry=false`).
Four
flusher startup callsites that must retry through transient "not ready"
errors
  were given explicit `RetryErr` escape hatches.

**D. segcore (C++→Go) error classification.**
A single shared Go-side table (`pkg/util/merr/segcore.go`) maps each
segcore code
to a merr sentinel + InputError/signal category, replacing scattered
hand-written
`if errorCode == ...` switches in the cgo wrappers. **Wire `Code` values
change
for every segcore pass-through error, not just the remapped ones.**
Named
sentinels remap (C++ `2003` → merr `2001`, `2033` → `2002`,
Folly/Knowhere codes
likewise); **all remaining pass-through codes (`2004`–`2043`, previously
surfaced to clients as raw C++ enum values) now serialize as `2000`**
(`ErrSegcore`), with the original C++ code preserved in the `Reason`
text
(`segcoreCode=...`); unknown/future codes collapse to `2000` as well
(pinned by
the `wire_code_projection` test). Transient segcore classes (object
storage /
file IO / OOM / mmap / FieldNotLoaded — 11 codes) now report
`Retriable=true`.
**Any client switching on raw segcore codes in the `2004`–`2043` range
must be
re-checked**; the in-Reason code remains available for diagnostics.
Signal
codes (PretendFinished / FollyCancel) are recognized centrally.
`errors.Is`-based
control flow on these (e.g. scheduler skip/retry) is preserved.

**E. InputError classification (25 sentinels + dynamic marks).**
25 sentinels in `errors.go` carry `WithErrorType(InputError)` (the
Collection /
ResourceGroup / Database families, `ErrIndexDuplicate`,
`ErrParameterInvalid`,
`ErrPrivilegeNotAuthenticated`, `ErrImportFailed`, `ErrQueryPlan`, ...),
plus dynamic
marks for the 8 segcore input codes (ExprInvalid, DimNotMatch,
MetricTypeInvalid, FieldIDInvalid, ...) and
`WrapErrAsInputError`. The widest blast radius is `ErrParameterInvalid`
(1100):
~2335 `WrapErrParameterInvalid*` callsites now classify as input /
non-retriable. Because of C1/C2 this changes retriability for
any path that returns these. **The audit to confirm no transient path
was
mis-marked is the single most important review item** (see Part 3). One
reverse
correction: storage field-stats parsing moved from `ErrParameterInvalid`
(input)
to `ErrDataIntegrity` — a corrupted stored stat is data corruption, not
user
input.

### Part 3 — Known risks & traps (called out proactively)

1. **`merr.Wrap` vs `WrapErr*Err` (code-masking).** `WrapErr*Err` builds
a
`wrappedMilvusError{sentinel: ErrServiceInternal}` whose `code()`
returns the
*outer* sentinel — it overwrites the inner typed code and hides the
`errors.Is`
chain. This is intentional (use it to *deliberately* downgrade), but it
was a
recurring conversion defect; the rule "add context with `merr.Wrap`,
downgrade
with `WrapErr*Err`" is enforced by convention and reviewed across the
diff.
2. **InputError × `retry.Do` blast radius.** Marking a sentinel
`InputError` makes
any `retry.Do(...)` without a `RetryErr` predicate stop retrying it.
Reviewers
should sanity-check that no transient use of the 19 newly-marked
sentinels
(especially `ErrParameterInvalid`) sits inside a retry loop that needed
to keep
   spinning. The known flusher cases were handled (see C2).
3. **The ~340 raw-error allowlist.** What remains as bare `errors.New`
is, by
design: package-level sentinels (caught by `errors.Is`), `//go:build
test`
sites, and out-of-band trees (`cmd/`, `tests/`, codegen, walimpls). The
linter
only bans the *direct-return* form; assignment-then-return escapes and
the full
no-exceptions ban are deferred to an AST-based linter (Tier 2,
documented).
4. **segcore C++ second step deferred.** This PR unifies classification
on the Go
side; splitting the dual-semantic C++ codes at the source is a
follow-up.

---

## Validation

- `make verifiers`: Go side clean (gofmt + static-check across modules,
including
  the new `rawmerrerror` rule with a 0-hit baseline repo-wide).
- `make test-go`: passing; the one real regression introduced (a
datanode
`invalid_task_type` assertion shifting `1` → `5` from a ParameterInvalid
  conversion) was fixed in-tree.
- go_client e2e CreateIndex assertions aligned to the new merr messages.

---------

Signed-off-by: zhenshan.cao <zhenshan.cao@zilliz.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
2026-06-12 15:04:51 -07:00

392 lines
11 KiB
Go

// Licensed to the LF AI & Data foundation under one
// or more contributor license agreements. See the NOTICE file
// distributed with this work for additional information
// regarding copyright ownership. The ASF licenses this file
// to you under the Apache License, Version 2.0 (the
// "License"); you may not use this file except in compliance
// with the License. You may obtain a copy of the License at
//
// http://www.apache.org/licenses/LICENSE-2.0
//
// Unless required by applicable law or agreed to in writing, software
// distributed under the License is distributed on an "AS IS" BASIS,
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
// See the License for the specific language governing permissions and
// limitations under the License.
package datacoord
import (
"context"
"time"
"github.com/hashicorp/golang-lru/v2/expirable"
"github.com/samber/lo"
"golang.org/x/exp/maps"
"github.com/milvus-io/milvus/internal/datacoord/allocator"
"github.com/milvus-io/milvus/internal/json"
"github.com/milvus-io/milvus/internal/metastore"
"github.com/milvus-io/milvus/pkg/v3/proto/internalpb"
"github.com/milvus-io/milvus/pkg/v3/taskcommon"
"github.com/milvus-io/milvus/pkg/v3/util/lock"
"github.com/milvus-io/milvus/pkg/v3/util/merr"
"github.com/milvus-io/milvus/pkg/v3/util/timerecord"
)
type ImportMeta interface {
AddJob(ctx context.Context, job ImportJob) error
UpdateJob(ctx context.Context, jobID int64, actions ...UpdateJobAction) error
GetJob(ctx context.Context, jobID int64) ImportJob
GetJobBy(ctx context.Context, filters ...ImportJobFilter) []ImportJob
CountJobBy(ctx context.Context, filters ...ImportJobFilter) int
RemoveJob(ctx context.Context, jobID int64) error
HandleCommitVchannel(ctx context.Context, jobID int64, vchannel string, callback func() error) error
AddTask(ctx context.Context, task ImportTask) error
UpdateTask(ctx context.Context, taskID int64, actions ...UpdateAction) error
GetTask(ctx context.Context, taskID int64) ImportTask
GetTaskBy(ctx context.Context, filters ...ImportTaskFilter) []ImportTask
RemoveTask(ctx context.Context, taskID int64) error
TaskStatsJSON(ctx context.Context) string
}
type importTasks struct {
tasks map[int64]ImportTask
taskStats *expirable.LRU[int64, ImportTask]
}
func newImportTasks() *importTasks {
return &importTasks{
tasks: make(map[int64]ImportTask),
taskStats: expirable.NewLRU[UniqueID, ImportTask](512, nil, time.Minute*30),
}
}
func (t *importTasks) get(taskID int64) ImportTask {
ret, ok := t.tasks[taskID]
if !ok {
return nil
}
return ret
}
func (t *importTasks) add(task ImportTask) {
t.tasks[task.GetTaskID()] = task
t.taskStats.Add(task.GetTaskID(), task)
}
func (t *importTasks) remove(taskID int64) {
task, ok := t.tasks[taskID]
if ok {
delete(t.tasks, taskID)
t.taskStats.Add(task.GetTaskID(), task)
}
}
func (t *importTasks) listTasks() []ImportTask {
return maps.Values(t.tasks)
}
func (t *importTasks) listTaskStats() []ImportTask {
return t.taskStats.Values()
}
type importMeta struct {
mu lock.RWMutex // guards jobs and tasks
jobs map[int64]ImportJob
tasks *importTasks
catalog metastore.DataCoordCatalog
}
func NewImportMeta(ctx context.Context, catalog metastore.DataCoordCatalog, alloc allocator.Allocator, meta *meta) (ImportMeta, error) {
restoredPreImportTasks, err := catalog.ListPreImportTasks(ctx)
if err != nil {
return nil, err
}
restoredImportTasks, err := catalog.ListImportTasks(ctx)
if err != nil {
return nil, err
}
restoredJobs, err := catalog.ListImportJobs(ctx)
if err != nil {
return nil, err
}
tasks := newImportTasks()
importMeta := &importMeta{}
for _, task := range restoredPreImportTasks {
t := &preImportTask{
importMeta: importMeta,
tr: timerecord.NewTimeRecorder("preimport task"),
times: taskcommon.NewTimes(),
}
t.task.Store(task)
tasks.add(t)
}
for _, task := range restoredImportTasks {
t := &importTask{
alloc: alloc,
meta: meta,
importMeta: importMeta,
tr: timerecord.NewTimeRecorder("import task"),
times: taskcommon.NewTimes(),
}
t.task.Store(task)
tasks.add(t)
}
jobs := make(map[int64]ImportJob)
for _, job := range restoredJobs {
jobs[job.GetJobID()] = &importJob{
ImportJob: job,
tr: timerecord.NewTimeRecorder("import job"),
}
}
importMeta.jobs = jobs
importMeta.tasks = tasks
importMeta.catalog = catalog
return importMeta, nil
}
func (m *importMeta) AddJob(ctx context.Context, job ImportJob) error {
m.mu.Lock()
defer m.mu.Unlock()
originJob := m.jobs[job.GetJobID()]
if originJob != nil {
originJob := originJob.Clone()
internalJob := originJob.(*importJob).ImportJob
internalJob.ReadyVchannels = lo.Union(originJob.GetReadyVchannels(), job.GetReadyVchannels())
job = originJob
}
err := m.catalog.SaveImportJob(ctx, job.(*importJob).ImportJob)
if err != nil {
return err
}
m.jobs[job.GetJobID()] = job
return nil
}
func (m *importMeta) UpdateJob(ctx context.Context, jobID int64, actions ...UpdateJobAction) error {
m.mu.Lock()
defer m.mu.Unlock()
if job, ok := m.jobs[jobID]; ok {
if job.GetState() == internalpb.ImportJobState_Completed ||
job.GetState() == internalpb.ImportJobState_Failed {
// import job is already completed or failed, no need to update
return nil
}
updatedJob := job.Clone()
for _, action := range actions {
action(updatedJob)
}
err := m.catalog.SaveImportJob(ctx, updatedJob.(*importJob).ImportJob)
if err != nil {
return err
}
m.jobs[updatedJob.GetJobID()] = updatedJob
}
return nil
}
func (m *importMeta) GetJob(ctx context.Context, jobID int64) ImportJob {
m.mu.RLock()
defer m.mu.RUnlock()
return m.jobs[jobID]
}
func (m *importMeta) GetJobBy(ctx context.Context, filters ...ImportJobFilter) []ImportJob {
m.mu.RLock()
defer m.mu.RUnlock()
return m.getJobBy(filters...)
}
func (m *importMeta) getJobBy(filters ...ImportJobFilter) []ImportJob {
ret := make([]ImportJob, 0)
OUTER:
for _, job := range m.jobs {
for _, f := range filters {
if !f(job) {
continue OUTER
}
}
ret = append(ret, job)
}
return ret
}
func (m *importMeta) CountJobBy(ctx context.Context, filters ...ImportJobFilter) int {
m.mu.RLock()
defer m.mu.RUnlock()
return len(m.getJobBy(filters...))
}
func (m *importMeta) RemoveJob(ctx context.Context, jobID int64) error {
m.mu.Lock()
defer m.mu.Unlock()
if _, ok := m.jobs[jobID]; ok {
err := m.catalog.DropImportJob(ctx, jobID)
if err != nil {
return err
}
delete(m.jobs, jobID)
}
return nil
}
func (m *importMeta) AddTask(ctx context.Context, task ImportTask) error {
m.mu.Lock()
defer m.mu.Unlock()
switch task.GetType() {
case PreImportTaskType:
err := m.catalog.SavePreImportTask(ctx, task.(*preImportTask).task.Load())
if err != nil {
return err
}
m.tasks.add(task)
case ImportTaskType:
err := m.catalog.SaveImportTask(ctx, task.(*importTask).task.Load())
if err != nil {
return err
}
m.tasks.add(task)
}
return nil
}
func (m *importMeta) UpdateTask(ctx context.Context, taskID int64, actions ...UpdateAction) error {
m.mu.Lock()
defer m.mu.Unlock()
if task := m.tasks.get(taskID); task != nil {
updatedTask := task.Clone()
for _, action := range actions {
action(updatedTask)
}
switch updatedTask.GetType() {
case PreImportTaskType:
err := m.catalog.SavePreImportTask(ctx, updatedTask.(*preImportTask).task.Load())
if err != nil {
return err
}
// update memory task
task.(*preImportTask).task.Store(updatedTask.(*preImportTask).task.Load())
case ImportTaskType:
err := m.catalog.SaveImportTask(ctx, updatedTask.(*importTask).task.Load())
if err != nil {
return err
}
// update memory task
task.(*importTask).task.Store(updatedTask.(*importTask).task.Load())
}
}
return nil
}
func (m *importMeta) GetTask(ctx context.Context, taskID int64) ImportTask {
m.mu.RLock()
defer m.mu.RUnlock()
return m.tasks.get(taskID)
}
func (m *importMeta) GetTaskBy(ctx context.Context, filters ...ImportTaskFilter) []ImportTask {
m.mu.RLock()
defer m.mu.RUnlock()
ret := make([]ImportTask, 0)
OUTER:
for _, task := range m.tasks.listTasks() {
for _, f := range filters {
if !f(task) {
continue OUTER
}
}
ret = append(ret, task)
}
return ret
}
func (m *importMeta) RemoveTask(ctx context.Context, taskID int64) error {
m.mu.Lock()
defer m.mu.Unlock()
if task := m.tasks.get(taskID); task != nil {
switch task.GetType() {
case PreImportTaskType:
err := m.catalog.DropPreImportTask(ctx, taskID)
if err != nil {
return err
}
case ImportTaskType:
err := m.catalog.DropImportTask(ctx, taskID)
if err != nil {
return err
}
}
m.tasks.remove(taskID)
}
return nil
}
func (m *importMeta) TaskStatsJSON(ctx context.Context) string {
tasks := m.tasks.listTaskStats()
ret, err := json.Marshal(tasks)
if err != nil {
return ""
}
return string(ret)
}
func (m *importMeta) HandleCommitVchannel(ctx context.Context, jobID int64, vchannel string, callback func() error) error {
m.mu.Lock()
defer m.mu.Unlock()
job := m.jobs[jobID]
if job == nil {
return merr.WrapErrImportSysFailedMsg("job %d not found", jobID)
}
// Idempotency: if vchannel already committed, skip.
for _, c := range job.GetCommittedVchannels() {
if c == vchannel {
return nil
}
}
if job.GetState() == internalpb.ImportJobState_Uncommitted {
updatedJob := job.Clone()
updatedJob.(*importJob).State = internalpb.ImportJobState_Committing
if err := m.catalog.SaveImportJob(ctx, updatedJob.(*importJob).ImportJob); err != nil {
return err
}
m.jobs[jobID] = updatedJob
job = updatedJob
}
// Move the job into commit phase before making any segment visible, then
// execute the callback before persisting the committed vchannel.
// If callback fails, we return error without persisting committed_vchannels;
// the caller retries and the callback will be invoked again. This avoids the
// scenario where committed_vchannels is persisted but callback fails, causing
// the idempotency check to skip the callback on retry (data stays invisible
// forever).
// The callback (setting is_importing=false) is idempotent, so re-execution on
// retry after a persist failure is safe.
//
// Visibility ordering note: the callback clears segment meta (is_importing=false)
// before this function persists job meta (committed_vchannels). Therefore a
// vchannel's imported data can become visible before the job-level transition
// to Completed (which happens later in checkCommittingJob once all vchannels
// have been recorded here). This is inherent to per-vchannel commit fences —
// 2PC for import is per-vchannel-atomic, not job-atomic. See MEP
// (milvus-io/milvus-design-docs#29) "Segment Visibility" section.
if err := callback(); err != nil {
return err
}
updatedJob := job.Clone()
updatedJob.(*importJob).CommittedVchannels = append(updatedJob.GetCommittedVchannels(), vchannel)
if err := m.catalog.SaveImportJob(ctx, updatedJob.(*importJob).ImportJob); err != nil {
return err
}
m.jobs[jobID] = updatedJob
return nil
}