* support cuda virtual devices
* disable NCCL path when virtual devices are used
* label virtual devices in description; add GPUx2 server CI jobs
* code refactor
PR #16308 set info.devices[id].integrated = false unconditionally for all
CUDA/HIP devices as a workaround for corrupted output on Jetson Orin
(#15034). On HIP/ROCm the device's real hipDeviceProp_t.integrated flag is
needed: with the cached field forced to false, supports_buft() refuses
CUDA host buffers on AMD APU/UMA parts, while get_type() already reads
prop.integrated (#23007) — an inconsistency that breaks integrated-GPU
host-buffer use on ROCm.
Guard the workaround so it only applies to non-HIP (CUDA) builds and
restore prop.integrated for HIP, keeping the Jetson workaround intact for
CUDA.
Fixes#23977
Signed-off-by: liminfei-amd <91481003+liminfei-amd@users.noreply.github.com>
* CUDA: dedup MoE gate/up activation quantization (fp4)
For MoE gate/up projections the src1 activation is broadcast across the
routed experts (ne11 == 1), so ids_src1 maps every one of a token's
n_expert_used slots to the same physical row. The MMQ path therefore
re-quantized each token's activation n_expert_used times.
For fp4 (NVFP4/MXFP4) src0, quantize each unique token row once instead of
once per expert. For NVFP4 a single quantize+scatter kernel
(quantize_scatter_mmq_nvfp4) quantizes each token once and writes the
resulting block_fp4_mmq straight to all n_expert_used slots, using an
inverse token->compact-row map (build_tok2c). MXFP4, and
GGML_CUDA_MOE_QUANT_GATHER=1, use a two-kernel variant: quantize unique
rows then gather into the expert-sorted layout (gather_mmq_fp4_blocks).
Both are bit-identical to the previous gather-then-quantize path (identical
source data, deterministic per-block quantization), verified by
test-backend-ops MUL_MAT_ID (type_a=nvfp4, broadcast b=1; 790/790 for the
default, gather, and per-expert paths) and by coherent end-to-end
generation. Set GGML_CUDA_NO_MOE_QUANT_DEDUP=1 to force the original
per-expert path.
Same-binary A/B on RTX 5090 (sm_120), Qwen3.6-35B-A3B-NVFP4 prefill @8192
(nsys, graphs-off; the unchanged mul_mat_q GEMM confirms stable clocks):
activation-quant GPU-busy drops 61% (78.2 -> 30.4 ms) with the fused
quantize+scatter, vs 33% (78.2 -> 52.8 ms) for the two-kernel gather. The
fused path avoids materializing and re-reading the 8x compact buffer,
writing the expert copies directly from registers.
* CUDA: bounds-check token ids in build_tok2c_kernel
Guard against malformed ids_src1: skip out-of-range token ids (t < 0 or
t >= n_tokens) and drop entries beyond n_expert_used per token instead of
writing past the token's tok2c region. No behavior change for valid MoE
routing data; test-backend-ops MUL_MAT_ID 790/790.
* Refactor the code based on review comments
- Removed previously added kernels that were not necessary anymore\
- Added an inverse mapping from (token, slot) to compact row. Each token is quantized once and scattered to its compact rows.
* Adding q8_1 support for dedup and addressing review comments
* Add pragma unrolls
* Remove redundant cudaMemsetAsync call
* Removing follow up redundancies
---------
Co-authored-by: praneshgo <227579474+praneshgo@users.noreply.github.com>
* opencl: workaround for A850 compiler compat
* opencl: fix DX compiler version parsing and cleanup
---------
Co-authored-by: Li He <lih@qti.qualcomm.com>
* opencl: exclude Adreno A7x from using Adreno MoE kernels
Some compilers for A7x devices miscompile the repack kernels, corrupting
the weights and causing MoE models to generate garbage output
* opencl: exclude A6x and unknown Adreno from MoE weights repack
* feat: Add shimmer text animation for processing state indicators
* feat: Redesign CollapsibleContentBlock component with improved UX
* feat: Add conditional setting display support with dependsOn field
* feat: Add showAgenticTurnStats setting for per-turn statistics
* feat: Update ChatMessageAgenticContent with improved UI and new features
* feat: Enhance file read tool UI/UX
* feat: Refine styling of collapsible content and code preview blocks
* feat: add terminal variant to CollapsibleContentBlock
* feat: add built-in tools UI registry
* feat: extract ChatMessageReasoningBlock and ChatMessageToolCallBlock
* refactor: simplify ChatMessageAgenticContent to use extracted blocks
* fix: correct markdown content block margin spacing
* fix: reorganize SettingsChatFields layout and reset button positioning
* fix: use direct map access in agentic store session methods
* refactor: remove reasoning preview/throttle system from CollapsibleContentBlock
* feat: add auto-scroll to reasoning block and remove showThoughtInProgress
* feat: add ChatMessageToolCallDateTime component and support for new tool types
* feat: improve auto-scroll reliability in reasoning block with RAF coalescing and MutationObserver
* feat: show MCP server favicon for tools without a built-in icon
* feat: add search-results parsing utilities and tests
* feat: add ChatMessageToolCallSearchResults component
* feat: integrate search results rendering into ChatMessageAgenticContent
* feat: display tool call input alongside output in ChatMessageToolCallBlock
* style: use muted foreground color in reasoning block content
* chore: Format
* feat: Refine reasoning block layout and make pending thoughts display configurable
* feat: Stream tool call code blocks with auto-scroll and handle partial JSON
* feat: add streaming permission gate infrastructure
* feat: wire permission gate into the agentic loop
* fix: bail out on abort and skip already-approved tool calls
* fix: clear partial tool calls on abort and savePartialResponse
* test: cover partial tool call cleanup end-to-end
* refactor: Remove streaming permission gate logic
* fix: Correct autoscroll and streaming gates for tool calls and reasoning blocks
* refactor: Chat Message Assistant componentization
* fix: Show health metadata for disabled MCP servers and promote connections on enable
* fix: Inherit global enabled state for missing MCP per-chat overrides
* refactor: Cleanup
* refactor: Split ChatMessageToolCallBlock into dedicated components
* feat: Add live streaming and auto-scroll for tool execution output
* feat: Add line numbers and change markers to file edit diffs
* chore: Formatting
* feat: Add type definitions and utilities for recommended MCP servers
* feat: Add recommended MCP servers configuration and storage key
* feat: Add McpServerCardCompact component for recommended servers
* feat: Add recommended servers section to Add New Server dialog
* feat: Update McpServerForm to support authorization requirements
* feat: Add select-none classes for text selection prevention
* feat: Add recommended MCP server icon assets
* refactor: Store dismissed MCP recommendations as a boolean flag
* feat: Render tool results as JSON or Markdown based on detected content type
* feat: UI improvement
* feat: Render search block early and update heading to show execution state
* fix: Prevent non-web-search tools from triggering the search UI block
* refactor: Cleanup
* refactor: Extract hardcoded icon size classes into shared constants
* refactor: Extract hardcoded tool result separator into a shared constant
* refactor: Tool Calls UI/logic
* refactor: Cleanup
* refactor: Cleanup
* refactor: Cleanup
* cuda : CUDA GGML_OP_LIGHTNING_INDEXER implementation (generic vector kernel + wmma kernel)
* chore : remove indentation of #pragma unroll
* cuda : remove unnecessary kernel template declarations
* cuda : add WARPS_PER_BLOCK and K_VECS_PER_BLOCK template parameters in lightning indexer kernels to avoid duplication of constants.
* cuda : relax MMA architecture requirements to Turing in lightning indexer implementation
* chore : renamed variables
* chore : rename ggml_cuda_op_lightning_indexer() to ggml_cuda_lightning_indexer()
* chore : TODO for AMD rocWMMA
* chore : whitespace formatting
* chore : another variable rename to fix problems caused by shadowing
* chore : yet another rename, this time uppercased all constants
* cuda : added alignment checks for Q and K tensors in lightning indexer implementation
---------
Co-authored-by: Stanisław Szymczyk <sszymczy@gmail.com>
* opencl: route `sub_group_shuffle_xor` to qcom ext when KHR ext is unavailable
KHR `sub_group_shuffle_xor` is not defined by compiler when
`cl_qcom_subgroup_shuffle` is present, causing certain FA
kernels fail to build. Define the KHR shuffle_xor using
the qcom extension.
* opencl: skip FA kernels with mixed and quant types for A7x to avoid compiler crash
read_file with append_loc emits "{n}\u2192 {line}". The space after the
arrow is meant as a separator, but it is indistinguishable from real
indentation. Models strip "{n}\u2192" yet keep the space, so the old_text
passed to edit_file carries a phantom leading space and never matches
(normalize_for_fuzzy_match trims trailing whitespace only, never leading).
Drop the separator space so the arrow abuts content: stripping "{n}\u2192"
now yields the exact line with its real indentation preserved, and the
failure mode cannot occur by construction. Update the description example
to match the new format.
In MODEL mode, modelPropsCache is never populated: fetchModelProps
call sites are gated on router-only state (isRouterMode checks,
routerModels always empty), so supportsThinking always reads an
empty chat template once a model is auto-selected.
Read serverStore.props.chat_template directly in non-router mode,
since the global /props already describes the single loaded model.
* ci : add HF_TOKEN to self-hosted workflows
Pass the HF_TOKEN_CI repo secret as HF_TOKEN env var in the self-hosted
build and server workflows.
Fix the stale build.yml path reference.
Assisted-by: pi:llama.cpp/Qwen3.6-27B
* cont : add comment
---------
Co-authored-by: ggerganov <ggerganov@users.noreply.github.com>
* metal: fuse snake activation (mul, sin, sqr, mul, add)
Mirror the CUDA, Vulkan and CPU snake fusion: same matcher on the naive
5-op chain, same F32 contract on a and inv_b, same F32/F16/BF16 kernel
with F32 compute. Follows the Metal backend idioms: bf16 instantiation
gated behind GGML_METAL_HAS_BF16 and concurrency ranges checked on the
remaining chain nodes before encoding, as done by the bin fusion.
Covered by the existing backend-agnostic SNAKE_FUSE tests.
* metal: absorb snake fusion into ggml_metal_op_bin
Extract the matcher to ggml_metal_op_can_fuse_snake, mirroring the
Vulkan naming, and dispatch the fused path from ggml_metal_op_bin.
The encode loop switch is back to a single call per case.
Address review from ggerganov
* metal: fix indentation in ggml_metal_op_can_fuse_snake
Raise the threshold for minimum buffer size from 1 GiB to 4 GiB, based
on real-world experiments of overcommitting device memory with model
weights larger than available VRAM, for example Qwen3.5-35B-A3B-Q8
running on a B70.
Also add a debug message to better track USM system allocations.
Signed-off-by: Francois Dugast <francois.dugast@intel.com>
* [SYCL] F16 (default) Flash Attention with XMX engine via oneDNN graph API; Qwen3.6-27b-Q8_0 prefill speed up x1.21 at p=512 and x4.26 at p=80k
* [SYCL] Address review on FA oneDNN path. Result: llama-bench---pp512; 32% increase with fa1; llama-perplexity---0.11% difference; tested model: mradermacher/Meta-Llama-3.1-8B-Instruct-Q8_0.gguf
* PR-25222 revision v2: addressed audits
* [SYCL] flash-attn oneDNN SDPA KV F16 rev 3.0: add BMG gate + multi-device sync. Narrow the scrope of this PR to Battlemage only (bmg; Xe2). Other archs (e.g., alchemist) fall back to existing FA kernel. When device_count >1, apply stream -> wait_and_throw(), validated working path for multi-gpu sync fix by @maxious.
Co-authored-by: maxious <81432+maxious@users.noreply.github.com>
* updated comment on bmg gate, noted the issue
---------
Co-authored-by: scientist3 <scientist.3@users.noreply.github.com>
Co-authored-by: hmscider <hmscider@users.noreply.github.com>
Co-authored-by: maxious <81432+maxious@users.noreply.github.com>
The f16 GEMV kernels take a vectorized path for ne00 >= 128 that casts the row
pointers to half4 or float4. When the row stride is not aligned, the wide load
becomes misaligned. On devices that require natural alignment for vector loads,
the kernel reads garbage. This is the case Intel GPUs and the kernels produce
incorrect results there. Adreno happpens to be byte addressable and the kernels
happen to work.
* server : clear checkpoints upon prompt clear
* server : move the prompt state data to the server_prompt_cache
Assisted-by: pi:llama.cpp/Qwen3.6-27B
* server : handle batched slot being cleared