mirror of
https://github.com/milvus-io/milvus.git
synced 2026-07-21 10:15:43 +00:00
issue: #51253 > Rebased on master after #51254 merged. The `db_name`/`db_id` part of the original PR is now covered by #51254; this PR is reduced to the two remaining gaps on the same query-by-id path. ## Problem When `DescribeCollection` is called with **only a `collectionID`** — no collection name, no db name (e.g. the HTTP management API on `:9091`) — two gaps remain after #51254: 1. **Top-level `collection_name` is empty** on both the cached provider path and the `describeCollectionTask` path. The response is initialized from the (empty) `request.CollectionName` before the name is resolved, and never backfilled. 2. **The cached provider discards the caller-provided collection id.** After deriving the name from the id, it re-resolves the id from that name via `GetCollectionID(db, name)`. Id-only lookups are cached under the request db name — empty here — so with same-name collections in different databases, a concurrent cache refresh can rebind the entry under the `""` key and the request would silently describe the wrong collection (the derived name matches, the id does not). ## Fix - `service_provider.go`: - Backfill `resp.CollectionName` from the resolved cached schema when the request carried no name. Requests that pass a name (including aliases) still echo it unchanged. - Skip the name→id re-resolution when the caller already supplied the id. `GetCollectionInfo` already validates the cache entry against the id (`collInfo.collID != collectionID` → refresh by id), so the caller-provided id is authoritative end to end. Name+id requests keep the existing name-precedence behavior. - Resolve identifiers on local copies instead of rewriting `request.CollectionName`/`CollectionID` in place — the access log interceptor and the success metric labels serialize the request after the handler returns, and previously recorded the resolved values instead of what the client sent. - `task.go`: backfill `t.result.CollectionName` from the coordinator result on the non-cached path. - `meta_cache.go`: fix a stale comment on `ResolveCollectionAlias` — `update()` now always keys `collInfo` by the real collection name (aliases live in the separate alias map), the old comment described pre-refactor behavior. - Tests: - `TestCachedProxyServiceProvider_DescribeCollection_ByIDFillsNameAndUsesRequestID`: drives the id-only path; `GetCollectionID` is deliberately not mocked, so any regression back to re-resolving the id panics the test. - `TestDescribeCollectionTask_FillsNameFromResultWhenQueriedByID`: same assertion for the non-cached path. - Removed the now-unused `GetCollectionID` expectation from the test added in #51254 (that call no longer happens on the id-only path). ## Follow-up commit: entity-keyed meta cache with a cluster-wide id index Reviewing this path surfaced a deeper issue in the proxy meta cache, fixed in the second commit: - By-id lookups (`GetCollectionName`, `GetCollectionInfo` with an empty name) linearly scanned a whole db bucket under the read lock, on every request even on cache hits. - The bucket scanned/filled was keyed by the *request* db name — empty for id-only calls — so entries landed in a bogus `""` bucket: a collection also cached under its real db was duplicated, and same-name collections of different databases evicted each other from the single `""`-keyed slot. The cache is now entity-keyed instead of request-keyed: - Fills land under the collection's **actual database** (carried in the describe response); the `""` bucket no longer exists. - A cluster-wide `collectionID → entry` index serves by-id lookups in **O(1)** regardless of the request db — matching rootcoord, which resolves by-id describes straight from `collID2Meta` without consulting the db name. - Name lookups normalize an empty db name to `default`, mirroring rootcoord's backward-compat normalization (`meta_table.getCollectionByNameInternal`). This also closes a latent cross-database mis-hit: a name lookup with an empty db could previously return whichever same-name collection was last id-cached under the `""` bucket. - All removal paths maintain the index (pointer-identity-guarded unindexing); invalidation sweeps stay exhaustive. ## Verification - Ran locally against a current master core build: the full `TestMetaCache*`, `TestCachedProxyServiceProvider*`, `TestDescribeCollectionTask*`, alias and partition test sets all pass, plus the **full `internal/proxy` suite** (only the pre-existing etcd-dependent `TestProxyRpcLimit` fails locally — it needs a live etcd, unrelated to this change). - New tests: `TestCachedProxyServiceProvider_DescribeCollection_ByIDFillsNameAndUsesRequestID` (id-only path; also asserts the request is not rewritten), `TestDescribeCollectionTask_FillsNameFromResultWhenQueriedByID`, `TestMetaCache_ByIDIndexRealDBBucket` (same-name collections in two databases filled by id: no eviction, entries under real dbs, by-name reuses by-id fill, every removal path cleans the index), `TestMetaCache_EmptyDBNameSharesDefaultEntry`. 🤖 Generated with [Claude Code](https://claude.com/claude-code) --- ## Update: proxy meta cache rebuilt around an id-primary store (2 new commits) Following review discussions on invalidation cost and the alias-DDL races, this PR now also rebuilds the cache (issue: #51533, design doc: `docs/design-docs/design_docs/20260716-proxy-metacache-id-inverted-index.md` in this PR): 1. **`fix: forward the pre-alter alias target id in the AlterAlias expiration`** — rootcoord resolves the pre-alter target under its database lock and carries it in the new additive `AlterAliasMessageHeader.old_collection_id`; the ack callback emits a second expiration entry so proxies evict BOTH AlterAlias targets by id. Closes the race where a concurrent Describe re-points the proxy's alias resolution before the expiration arrives, leaving the old target permanently stale. Older rootcoords (field 0) fall back to the proxy's hint resolution. 2. **`enhance: rebuild the proxy meta cache around an id-primary store`** — the primary store is keyed by the cluster-unique collection id (single source of truth, with a per-database generation); name/alias resolution become hints validated against the primary on read; partition caches are keyed by id; fills are ordered against invalidations by a fill RWMutex (drain in-flight describes before evicting), replacing the per-collection timestamp floor. Every invalidation becomes O(1) — no more full-cache scans under the write lock on Load/Release/Drop/Rename/Alter broadcasts, and DropDatabase/AlterDatabase becomes a generation bump. Also removes a latent stale read (alias-keyed partition entries surviving DropCollection). Verification: targeted cache suites green (rename, alias re-point under concurrent describe, drop+recreate with a reused name, db-generation invalidation, cross-db isolation, gated-mock fill/invalidation ordering); three negative controls (hint validation / dbGen check / fill drain disabled) each fail exactly the test guarding them; full-package `-race` delegated to CI. --- ## Update: simplification series (final form) Following design review, the cache was iteratively simplified to an **id-primary store with declared hints** — every mechanism that could be replaced by an invariant was deleted (net-negative diffs throughout): - Primary store keyed by the cluster-unique collection id; liveness IS presence. Name/alias resolution are hints written ONLY together with their entry (declared aliases) and validated against the primary on every read — a stale hint can only cost one extra describe, never a stale read. - Fill/invalidation ordering via a fill RWMutex (drain in-flight describes before evicting); the per-collection timestamp floor, database generations, background GC, alias negative cache, per-partition cache, and the resolve-and-forget DescribeAlias path are all **deleted**. Evictions clean everything the entry owns synchronously; **no invalidation path performs any RPC**. - Old-rootcoord fallbacks: id-describes without DbName are served uncached; alias-DDL broadcasts without ids trigger hint resolution plus a holder scan **gated on that exact fingerprint** (dead code once rootcoords are upgraded), closing the ghost-alias case with no healing event. - **Accepted gaps (WONT-FIX)** are recorded in the design doc §6 — notably the upgrade-window new-target `Aliases` display lag (the association exists only at rootcoord; self-heals on first use; display-only). See `docs/design-docs/design_docs/20260716-proxy-metacache-id-inverted-index.md` for the full design, invariants and trade-offs. Signed-off-by: xiaofanluan <xf@hjjaq.com> Co-authored-by: xiaofanluan <xf@hjjaq.com> Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
151 lines
6.4 KiB
Go
151 lines
6.4 KiB
Go
// Licensed to the LF AI & Data foundation under one
|
|
// or more contributor license agreements. See the NOTICE file
|
|
// distributed with this work for additional information
|
|
// regarding copyright ownership. The ASF licenses this file
|
|
// to you under the Apache License, Version 2.0 (the
|
|
// "License"); you may not use this file except in compliance
|
|
// with the License. You may obtain a copy of the License at
|
|
//
|
|
// http://www.apache.org/licenses/LICENSE-2.0
|
|
//
|
|
// Unless required by applicable law or agreed to in writing, software
|
|
// distributed under the License is distributed on an "AS IS" BASIS,
|
|
// WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied.
|
|
// See the License for the specific language governing permissions and
|
|
// limitations under the License.
|
|
|
|
package rootcoord
|
|
|
|
import (
|
|
"context"
|
|
"testing"
|
|
|
|
"github.com/stretchr/testify/mock"
|
|
"github.com/stretchr/testify/require"
|
|
|
|
mockrootcoord "github.com/milvus-io/milvus/internal/rootcoord/mocks"
|
|
"github.com/milvus-io/milvus/internal/util/proxyutil"
|
|
"github.com/milvus-io/milvus/pkg/v3/proto/proxypb"
|
|
"github.com/milvus-io/milvus/pkg/v3/streaming/util/message"
|
|
"github.com/milvus-io/milvus/pkg/v3/util/funcutil"
|
|
)
|
|
|
|
// TestDDLCallbacksAlterAliasV2AckCallback_OldTargetRouting locks in how the
|
|
// AlterAlias ack callback routes the OLD target eviction by the header's
|
|
// OldCollectionId, which shares the message type with CreateAlias. The zero
|
|
// value is the compat-safe scan path so a replayed old-rootcoord AlterAlias
|
|
// (which predates old_collection_id and leaves it 0) still closes its ghost:
|
|
// - > 0 : old target known -> evict it by id (no scan).
|
|
// - == 0 : old target UNKNOWN (new AlterAlias could not resolve it, OR an old
|
|
// rootcoord that never set the field) -> emit a CollectionID==0 alias entry
|
|
// so the proxy holder-scans (graceful degrade, never fails the alter).
|
|
// - < 0 : CreateAlias sentinel (provably no old target) -> NO CollectionID==0
|
|
// entry, so a normal create never triggers the O(N) holder scan.
|
|
func TestDDLCallbacksAlterAliasV2AckCallback_OldTargetRouting(t *testing.T) {
|
|
ctx := context.Background()
|
|
const newID = int64(20)
|
|
invalidate := func(t *testing.T, oldCollectionID int64) []*proxypb.InvalidateCollMetaCacheRequest {
|
|
meta := mockrootcoord.NewIMetaTable(t)
|
|
meta.EXPECT().AlterAlias(mock.Anything, mock.Anything).Return(nil)
|
|
|
|
var got []*proxypb.InvalidateCollMetaCacheRequest
|
|
pcm := proxyutil.NewMockProxyClientManager(t)
|
|
pcm.EXPECT().InvalidateCollectionMetaCache(mock.Anything, mock.Anything, mock.Anything).RunAndReturn(
|
|
func(ctx context.Context, req *proxypb.InvalidateCollMetaCacheRequest, opts ...proxyutil.ExpireCacheOpt) error {
|
|
got = append(got, req)
|
|
return nil
|
|
}).Maybe()
|
|
|
|
c := newTestCore(withMeta(meta), withTsoAllocator(newMockTsoAllocator()))
|
|
c.proxyClientManager = pcm
|
|
cb := &DDLCallback{Core: c}
|
|
|
|
raw := message.NewAlterAliasMessageBuilderV2().
|
|
WithHeader(&message.AlterAliasMessageHeader{
|
|
DbName: "db",
|
|
Alias: "a",
|
|
CollectionId: newID,
|
|
OldCollectionId: oldCollectionID,
|
|
}).
|
|
WithBody(&message.AlterAliasMessageBody{}).
|
|
WithBroadcast([]string{funcutil.GetControlChannel("test")}).
|
|
MustBuildBroadcast()
|
|
msg := message.MustAsBroadcastAlterAliasMessageV2(raw)
|
|
require.NoError(t, cb.alterAliasV2AckCallback(ctx, message.BroadcastResultAlterAliasMessageV2{
|
|
Message: msg,
|
|
Results: map[string]*message.AppendResult{},
|
|
}))
|
|
return got
|
|
}
|
|
|
|
// helpers over the captured requests
|
|
hasIDOnly := func(reqs []*proxypb.InvalidateCollMetaCacheRequest, id int64) bool {
|
|
for _, r := range reqs {
|
|
if r.GetCollectionID() == id && r.GetCollectionName() == "" {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|
|
hasScanTrigger := func(reqs []*proxypb.InvalidateCollMetaCacheRequest) bool {
|
|
// the proxy holder-scan is gated on CollectionID==0 + the alias name
|
|
for _, r := range reqs {
|
|
if r.GetCollectionID() == 0 && r.GetCollectionName() == "a" {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|
|
// hasID matches any request carrying the id, regardless of an accompanying
|
|
// name (the new-target eviction rides on the first request, which also carries
|
|
// the alias name, so hasIDOnly deliberately does NOT match it).
|
|
hasID := func(reqs []*proxypb.InvalidateCollMetaCacheRequest, id int64) bool {
|
|
for _, r := range reqs {
|
|
if r.GetCollectionID() == id {
|
|
return true
|
|
}
|
|
}
|
|
return false
|
|
}
|
|
|
|
t.Run("old target known -> evict by id, no scan trigger", func(t *testing.T) {
|
|
reqs := invalidate(t, 10)
|
|
require.True(t, hasIDOnly(reqs, 10), "old target 10 must be evicted by id")
|
|
require.False(t, hasScanTrigger(reqs), "a known old target must not trigger the holder scan")
|
|
})
|
|
|
|
t.Run("old target unknown (0) -> scan trigger, no stray id eviction", func(t *testing.T) {
|
|
// zero covers both a new AlterAlias that could not resolve and a replayed
|
|
// old-rootcoord AlterAlias that never set old_collection_id.
|
|
reqs := invalidate(t, 0)
|
|
require.True(t, hasScanTrigger(reqs), "an unknown old target (0) must emit a CollectionID==0 scan trigger")
|
|
require.False(t, hasIDOnly(reqs, 0), "zero must never be sent as a real id eviction")
|
|
})
|
|
|
|
t.Run("create alias sentinel (-1) -> no scan trigger", func(t *testing.T) {
|
|
reqs := invalidate(t, aliasNoOldTarget)
|
|
require.False(t, hasScanTrigger(reqs), "CreateAlias (aliasNoOldTarget) must not trigger an O(N) holder scan")
|
|
})
|
|
|
|
t.Run("new target is always evicted by id", func(t *testing.T) {
|
|
// the new target's canonical entry must be expired regardless of the old
|
|
// target's fate; a regression dropping the unconditional first eviction
|
|
// would leave an id-only Describe of the new target serving a stale
|
|
// Aliases list.
|
|
for _, oldID := range []int64{10, 0, aliasNoOldTarget} {
|
|
reqs := invalidate(t, oldID)
|
|
require.True(t, hasID(reqs, newID), "new target %d must be evicted (oldID=%d)", newID, oldID)
|
|
}
|
|
})
|
|
|
|
t.Run("old target == new target -> deduped, no redundant eviction, no scan", func(t *testing.T) {
|
|
// the guard is OldCollectionId > 0 && OldCollectionId != CollectionId; when
|
|
// they coincide the old-target branch must be skipped so we do not emit a
|
|
// second, redundant eviction for the same id, and never fall into the scan.
|
|
reqs := invalidate(t, newID)
|
|
require.True(t, hasID(reqs, newID), "the new target must still be evicted")
|
|
require.False(t, hasIDOnly(reqs, newID), "no separate id-only old-target eviction when old==new")
|
|
require.False(t, hasScanTrigger(reqs), "a resolved (equal) old target must not trigger the scan")
|
|
})
|
|
}
|