## Summary
`NullVectorQueryChecker` was producing spurious failures during chaos
tests because it incorrectly treated all-null query results as "possible
data corruption".
**Root cause**: The checker collects nullable vector fields from
`describe_collection` at init time. When a persistent collection is
reused across test runs, it already contains `new_vec_XXX` fields added
by `AddVectorFieldChecker` in previous runs. Segments sealed *before*
those fields were added have all-null values for them by design — this
is correct behavior. A `limit=50` query with no filter can return 50
rows all from pre-existing segments → all null → false alarm. During
s3-pod-failure chaos, fewer new non-null rows were being flushed
successfully, making the all-null result more frequent and pushing the
succ rate below the pass threshold.
**Fix (two changes):**
1. Exclude `new_vec_*` fields from `nullable_vector_fields` at init —
dynamically-added fields are already verified by `AddVectorFieldChecker`
immediately after each insertion
2. Use a `{field} is not null` filter in the query — makes the check
semantically correct: verify that rows with non-null vector values
remain queryable under chaos, and that Milvus applies the filter
correctly
issue: #49262
## Test plan
- [ ] Chaos test `test_concurrent_operation.py` with s3-pod-failure —
`null_vector_query` succ rate should no longer drop below threshold
- [ ] `NullVectorQueryChecker` still catches genuine filter failures:
rows returned by `is not null` filter that contain null values are
correctly reported as failure
🤖 Generated with [Claude Code](https://claude.com/claude-code)
Signed-off-by: yanliang567 <82361606+yanliang567@users.noreply.github.com>
Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Tests
E2E Test
Configuration Requirements
Operating System
| Operating System | Version |
|---|---|
| Amazon Linux | 2023 or above |
| Ubuntu | 20.04 or above |
| Mac | 10.14 or above |
Hardware
| Hardware Type | Recommended Configuration |
|---|---|
| CPU | x86_64 architecture Intel CPU Sandy Bridge or above CPU Instruction Set - SSE4_2 - AVX - AVX2 - AVX512 or arm64 Linux/MacOS |
| Memory | 16 GB or more |
Software
| Software Name | Version |
|---|---|
| Docker | 19.05 or above |
| Docker Compose | 1.25.5 or above |
| jq | 1.3 or above |
| kubectl | 1.14 or above |
| helm | 3.0 or above |
| kind | 0.10.0 or above |
Installing Dependencies
Troubleshooting Docker and Docker Compose
- Confirm that Docker Daemon is running:
$ docker info
-
Ensure that Docker is installed. Refer to the official installation instructions for Docker CE/EE.
-
Start the Docker Daemon if it is not already started.
-
To run Docker without
rootprivileges, create a user group labeleddocker, then add a user to the group withsudo usermod -aG docker $USER. Log out and log back into the terminal for the changes to take effect. For more information, see the official Docker documentation for Managing Docker as a Non-Root User.
- Check the version of Docker-Compose
$ docker compose version
docker compose version 1.25.5, build 8a1c60f6
docker-py version: 4.1.0
CPython version: 3.7.5
OpenSSL version: OpenSSL 1.1.1f 31 Mar 2020
- To install Docker-Compose, see Install Docker Compose
Install jq
Install kubectl
Install helm
- Refer to https://helm.sh/docs/intro/install/
Install kind
Run E2E Tests
$ cd tests/scripts
$ ./e2e-k8s.sh
Getting help
You can get help with the following command:
$ ./e2e-k8s.sh --help
Python Code Quality (ruff via uv)
Ruff is configured at tests/pyproject.toml and covers all Python code under tests/
(python_client/, restful_client/, restful_client_v2/, benchmark/, scripts/).
uv is only used to host the lint/format toolchain; each
sub-directory continues to manage its runtime dependencies via its own requirements.txt.
$ cd tests/
$ uv sync # install ruff into a local .venv
$ uv run ruff check . # lint
$ uv run ruff check . --fix # lint with auto-fix
$ uv run ruff format . # format in place
$ uv run ruff format --check . # format check only (CI-friendly)
Rules enabled: E, F, W, I, UP. Target Python version: 3.10.
Lint only PR-changed files (Python Lint CI parity)
The GitHub Actions workflow .github/workflows/python-lint.yaml (job
"Python Lint (tests/) / Ruff (changed files only)") runs ruff check
and ruff format --check against the changed tests/**/*.py set on
every PR. Running uv run ruff check . over the whole tree is too coarse
because the historical contents of many files predate the lint config and
will fail unrelated rules.
To reproduce the CI step locally, use the Makefile shipped in this directory:
$ cd tests/
$ make ci # ruff check + format --check on PR-changed *.py (CI equivalent)
$ make lint-fix # ruff check --fix on PR-changed *.py
$ make format # ruff format on PR-changed *.py
$ make help # show all targets and the detected BASE_REF
BASE_REF is auto-detected from the current branch's open PR via gh pr view:
the PR URL is parsed to discover the base <owner>/<repo> and matched against
your local git remotes, producing e.g. upstream/master, upstream/2.x, or
origin/main for direct clones.
If no open PR exists for the branch, you must set BASE_REF explicitly
(no guessing — a wrong base diffs against unrelated commits):
$ make ci BASE_REF=upstream/master
Requires uv (provides uvx) and gh authenticated against GitHub
(gh auth status). The ruff version is pinned in the Makefile via
RUFF_VERSION to match the workflow.