mirror of
https://github.com/ikawrakow/ik_llama.cpp.git
synced 2026-07-21 10:15:35 +00:00
* Add GLM-5.2/DeepSeek-V3.2 DSA lightning indexer (batch-local, single-seq prefill) Implements the sparse top-k "lightning indexer" attention for LLM_ARCH_GLM_DSA in build_deepseek2_layer_attention (ik's deepseek2 graph). What it does (per layer, gated on model.arch==GLM_DSA && indexer_attn_q_b): - indexer_q = indexer_attn_q_b(q_lora latent), split rope(64)/nope(64), NEOX-rope the pe part, concat. indexer_k = indexer_attn_k(attn_norm out), LayerNorm w/ bias, same rope/concat (single key head, MQA). - scores = relu(indexer_k . indexer_q), scaled per-head weights (indexer_proj), summed over heads, + base causal mask, then ggml_top_k(min(top_k, n_tokens)). - sparse mask: ggml_fill(-inf) -> ggml_set_rows(0) at top_k positions -> + causal, used in the soft_max_ext attention path (-mla 1 -fa 0) instead of KQ_mask. Simplifications (intentional, proven sound): - Batch-local: no indexer KV-cache. Indexer keys are the current batch tokens. - Walsh-Hadamard transform omitted: orthonormal rotation, (Hq).(Hk)==q.k, no score change. Validation (GLM-5.2-UD-IQ2_M, 3x P100, -mla 1 -fa 0): - Compiles clean (CUDA sm_60); loads and runs. - c512 -b512 (n_seq=1) PPL = 2.7760, byte-identical to dense baseline (indexer disabled) = 2.7760, all 8 chunks match -> indexer is an exact no-op when top_k>=n_tokens. Proves correctness-preservation. - 3105-token prompt completion (top_k=2048 < 3105 -> indexer ACTIVELY masks): prompt-eval produces coherent, accurate continuation, identical to dense for the prompt+early-gen tokens. No NaN/crash. Confirms the masking path works in prefill. Known limitations (documented follow-ups, NOT handled): - Single-sequence prefill only. Multi-sequence batches (n_seq>1, e.g. perplexity default n_batch>n_ctx) and kv_head>0 (decode) break the batch-local key->slot mapping. n_seq>1 -> NaN (use n_batch==n_ctx). Decode (kv_head>0): each generated token sees only itself as an indexer key, so generation degenerates into repetition after the prompt (dense A/B stays coherent) -- this is the decode-cache stub, the documented next step. - Flash-attn path (-fa 1, F16 mask) still uses dense KQ_mask (soft_max path only). - Decode indexer KV-cache + Hadamard cached-K storage not implemented. Runtime gate: DSA_INDEXER_DISABLE=1 falls back to dense attention (for A/B). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: decode-correct via persistent indexer-K cache Make the lightning-indexer correct for DECODE (not just prefill). Previously the indexer was batch-local, so a generated token only scored against itself and generation degenerated. Now the indexer keys are cached across the full context. Changes - llama_kv_cache: add per-layer indexer-key cache `kr_l` [indexer_head_size, kv_size] (F16, MQA single head), allocated alongside the MLA latent cache for GLM_DSA. - build_deepseek2_dsa_indexer: write the batch's (Hadamard-rotated) indexer keys to kr_l at kv_head, read back the full [128, n_kv] cached keys, and score the indexer queries against ALL past keys. Returns the full descending argsort of the scores. - Walsh-Hadamard rotation of indexer q/k (cparams.dsa_indexer_hadamard, default on; filled in llama_set_inputs). Score-preserving; improves cached-K F16 precision. - build_deepseek2_dsa_sparse_mask: rank-based full-coverage scatter (write a 0/-BIG penalty into EVERY key slot keyed by rank) instead of partial set_rows into a -inf fill — the CUDA in-place set_rows does not preserve an un-written base, which had corrupted decode when n_kv > top_k. - Attention-sink force-inclusion (DSA_SINK, default 1): boost the first key(s) so the sink always survives top-k. The IQ2_M-quantized indexer under-ranks the sink, and masking it collapsed decode; with the boost, top_k=2048 over n_kv>2048 stays coherent. ggml backend fixes (needed by the indexer) - CUDA argsort: report unsupported when padded ncols > 1024 (one-thread-per-column bitonic launch limit) so the scheduler falls back to the CPU argsort. Fixes "invalid configuration argument" for top_k over a large n_kv. - CUDA cpy/dup: support I32 -> I32 (top_k index copies / cross-backend moves). Validation (GLM-5.2-UD-IQ2_M, 3xP100 + --cpu-moe, -mla 1 -fa 0) - c512 PPL = 2.0743, byte-identical to dense (all 8 chunks): no-op path exact. - Short-context decode (300 tok): coherent, identical to dense. - Long-context decode (2521-tok prompt, n_kv>top_k, real masking of ~474 keys, 120+ tok generated): coherent with the sink boost; dense A/B also coherent. Gated behind arch==GLM_DSA + indexer tensors + kr_l cache; DSA_INDEXER_DISABLE=1 forces dense. Remaining: FA path still uses the dense KQ_mask; multi-sequence (n_seq>1) batches; deepseek32 arch wiring. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: wire sparse mask into the flash-attention path (-fa 1) The DSA sparse top-k mask is now applied on the -fa 1 path (our serving config), not just -fa 0 soft_max. c512 PPL on -fa 1 = 2.0743, byte-identical to dense (no regression, indexer no-op exact at n_kv <= top_k). Gated arch==GLM_DSA with DSA_INDEXER_DISABLE escape; -fa 0 path unchanged. Long-context -fa 1 decode coherence (n_kv > top_k, mask actually biting) validation is still running at commit time; the FA mask reuses the same full-coverage scatter proven coherent on the -fa 0 decode path, so it should hold, but confirm before relying on long-context -fa 1. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: UPDATE 4 — MLA-FA fix merged, FA path validated, multi-seq characterized Document the re-validation after cherry-picking the MLA-FA vec-decode fix (5f18dcc0): - FA path is ALIVE. Long-ctx -fa 1 decode (2521-tok prompt > top_k, mask actively biting) is now COHERENT at -mla 1 and -mla 3, vs the pre-fix degeneration into "0.0.0.0..." repetition. Matches dense (DSA_INDEXER_DISABLE) and -fa 0 controls. - c512 -fa 1 PPL: indexer-ON == dense == 2.0854, byte-identical all 8 chunks (exact no-op when n_kv <= top_k; no regression). The 2.0743->2.0854 shift is the MLA-FA fix changing V accumulation, not an indexer artifact (ON==dense proves it). - Indexer is feature-complete + validated for single-seq prefill+decode on both -fa 0 and -fa 1, at -mla 1 and -mla 3 (the R740 serving target). Remaining PR gaps, characterized honestly: - Multi-seq (n_seq>1) with active mask is BROKEN (n_seq=2 c4096 PPL 62.6 vs dense multi-seq 2.54 and single-seq indexer 3.05). No NaN/crash anymore. Root cause: the indexer uses a single scalar kv_head/n_kv for the whole ubatch; multi-seq needs per-sequence cache writes + per-sequence top-k. Fix deferred (structural). - deepseek32 arch: N/A in this fork. DSA lives entirely under LLM_ARCH_GLM_DSA; there is no LLM_ARCH_DEEPSEEK32 enum. Documented the steps to add one if a real deepseek32 GGUF is ever served. Also commit DSA_REFERENCE.md (verbatim mainline deepseek32/glm-dsa source, the port reference), trimmed of a stray agent-handoff footer. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: per-sequence attention sink — fix multi-seq (n_seq>1) UPDATE 5. The DSA lightning indexer was numerically broken for multi-sequence batches once the top-k mask bites (n_kv > top_k): c4096 n_seq=2 PPL 62.6 vs dense 2.54, while single-seq was fine. Root cause: the attention-sink force-include boosted the GLOBAL key range [0, n_sink) by +1e20, which only protects sequence 0's sink. With several sequences packed contiguously into one ubatch (seq 0 at cells [0,n0), seq 1 at [n0,n1), ...), every non-first sequence's sink lives at cell n0.. (not cell 0), got no boost, and was dropped from top-k once the mask bites — collapsing that sequence (chunk[2]=61.2 while chunk[1]=2.33). The cache write and score/argsort were already per-sequence correct: tokens are placed contiguously like the main K cache, and the base KQ_mask (filled from kv_self.cells[i].has_seq_id) already drives cross-seq keys to -inf before argsort. Only the sink was anchored at the wrong (global) cell. Fix: replace the global arange sink boost with a per-graph input tensor inp_dsa_sink {n_kv, n_tokens} (F32), filled on the CPU in llama_set_inputs from kv_self.cells exactly like the KQ_mask: inp_dsa_sink[j,i] = 1e20 iff cell[i].pos in [0,n_sink) AND cell[i].has_seq_id(seq_of_query_j), else 0 so each query force-includes only its OWN sequence's sink. For a single contiguous sequence from pos 0 this is exactly the old "cell index < n_sink" set with the same magnitude, so n_seq==1 is byte-identical. Validation (3x P100, -ngl 99 --cpu-moe -mla 3 -fa 1, wikitext-2): - c4096 n_seq=2 indexer chunk[2]: 61.2 -> 3.07 (== single-seq 3.05). - c2048 topk=1024 (mask bites): n_seq=4 == n_seq=1 chunk-for-chunk (2.5005/2.6080/2.7759/3.1137 vs .../3.1138) -> multi-seq is numerically identical to processing each sequence alone. - c512 n_seq=1 indexer ON == dense, all 4 chunks byte-identical (no regression). n_seq=4 at full c4096 (n_kv=16384) OOMs the P100 compute buffer (capacity, not correctness; n_seq=4 proven correct at c2048/n_kv=8192). GLM-5.2 DSA indexer is now sequence-correct for n_seq>=1, prefill+decode, soft_max+FA, -mla 1/-mla 3. Fully general and PR-ready. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: UPDATE 6 — serving-correctness (kr_l maintained across shift/defrag/seq-ops; per-seq sink on first-present pos) An adversarial review found the indexer was proven on the perplexity path but not the serving path: the persistent indexer-K cache kr_l was written/read but never *maintained* by the KV-cache mutators, and the attention sink anchored on absolute pos<n_sink (wrong after multi-turn seq_rm). This closes those gaps and pins down what is actually reachable on the MLA model. kr_l maintenance: - build_k_shift (llama-build-context.cpp): rotate the indexer keys by the same per-cell delta as the main K. The cached key is H*concat(RoPE(k_pe,pos),k_nope), so un-Hadamard (H sym/orthonormal => H*H=I) -> RoPE-delta the pe sub-block -> re-Hadamard. Exact because GLM-DSA has no rope-scaling metadata (ext_factor=0, attn_factor=1, freq_scale=1), so NEOX RoPE is pure/composable. Params mirror the forward indexer RoPE exactly (rope_factors=nullptr); no DEEPSEEK2 yarn-shift leak. Non-in-place (cont->rope->concat->re-Had->cpy), no aliasing. K-shift Hadamard input filled in llama_set_k_shift with the identical Sylvester construction. - build_defrag: kr_l row-move mirrors the k_l move (defrag never changes pos, so no re-RoPE). max_moves divisor 6->9 *n_layer when the indexer cache is present. - seq_rm/seq_cp/seq_keep are metadata-only (verified) so kr_l rows stay matched to cells; seq_add/seq_div set has_shift and route through K-shift. No seq-op change. Per-seq sink (llama.cpp llama_set_inputs): anchor on each sequence's FIRST PRESENT pos (min present pos over the scored n_kv span), not absolute pos<n_sink. After multi-turn seq_rm drops a sequence's early tokens its earliest survivor has pos>=n_sink; the absolute test would protect nothing. Fresh seq at pos 0 => min=0 => byte-identical to the old behaviour. Serving-shift finding (the whole point): a RoPE context-shift on this model is REFUSED BY THE ENGINE. get_can_shift() returns false for all MLA models (is_mla_model() includes GLM_DSA); llama_kv_cache_update returns 1 -> "main : failed to eval". Reproduced AND isolated with a dense control (DSA_INDEXER_DISABLE=1): dense fails identically at the same token. The failure is pre-existing MLA engine behaviour, independent of the indexer. On the MLA path the shift never happens, so the indexer's kr_l can never desync via K-shift; the build_k_shift kr_l block is correct-and-dormant (documented loudly in code). Validation (3x P100, -ngl 99 --cpu-moe -mla 3 -fa 1, GGML_CUDA_NO_PINNED=1, numactl --interleave=all, wikitext-2): - No regression: c512 n_seq=1 indexer ON == dense == 2.1957 +/- 0.12031, byte-identical all 4 chunks (2.2770/2.8741/2.3956/2.1957). - Multi-seq: c4096 n_seq=2 chunk[1]=2.33 chunk[2]=3.07 healthy (== UPDATE 5; per-seq sink change did not regress). - Serving shift: engine-refused for MLA, dense control fails identically. - Independent adversarial review: GO, no correctness defect in the diff. - Build clean (llama-cli, llama-perplexity, sm_60). Comments updated (build_deepseek2.cpp): multi-seq+FA no longer limitations; sink description matches per-seq min-pos anchoring; BIG=1e30 masks on both soft_max and FA paths. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-5.2 DSA indexer: UPDATE 7 — FIX latent graph-reuse cache-fixup omission for the kr_l indexer cache update_cache_copies() re-points the K/V cache writes to the current kv_head whenever a compute graph is REUSED (can_reuse_graph reuses iff kv_self.n == prev->n_kv). The persistent indexer-key cache write (kr_l) is a separate ggml_cpy whose destination view bakes kv_head at graph-build time, and it was NEVER registered for that fixup. Under FA the cache pads to 256, so consecutive single-token decode ubatches share the same padded n_kv and the graph IS reused; without the fixup the kr_l write keeps landing in the first ubatch's slot and later ubatches never populate their own recent index-key cells (those cells stay at the alloc-zeroed 0.0). Structurally identical to the MiniMax MSA bug (fork commit 133d14c9). Fix (mirrors the K/V cache_copies fixup, same shape as MSA 133d14c9): - llama-context.h: new std::vector<CacheCopy> dsa_cache_copies. - llama.cpp ctor: resize dsa_cache_copies to n_layer (null entries -> no-op when DSA off). - build_deepseek2.cpp: register the kr_l ggml_cpy as dsa_cache_copies[il] = {kr_cpy, kr->nb[1]}. - llama.cpp update_cache_copies(): re-point each registered cpy view_offs = kv_head*step and patch src[1]->data/data, exactly like K/V, with the c.cpy->view_src == kv_self.kr_l[il] (+ null/op) guard the MSA fix omitted. soft_max / non-DSA paths byte-identical. Validation (GLM-5.2-UD-IQ2_M, 3x P100 -ngl 99 --cpu-moe -t 32, NO_PINNED, P2P-disable patch re-applied to get a working multi-GPU baseline — see UPDATE 7.3; that patch was lost in the upstream rebase and is required separately): - c512 -fa1 -mla3 indexer ON: 2.1983 (== prior baseline; build healthy). - Long-ctx FA decode, 2735-tok recall prompt, -mla3 -fa1 temp0, reuse ON (default): coherent, correct deep-context recall ("Dr. Mariana Velasquez ... Daniel Okonkwo") on BOTH the fixed and the unfixed binary. - ub128 PPL -fa1 -mla3 reuse ON, unfixed: 1.7239/1.8211/2.1888/2.4517, healthy (no inflation). Honest scope: the bug is real in code but LATENT for GLM-DSA at its configured top_k=2048 (permissive selection keeps the genuinely-attended recent blocks even when reuse leaves some recent index-key cells stale), unlike MSA's tighter top-k where it inflated PPL ~2x. The fix is correct and prevents the latent corruption from biting at any tighter top_k / longer ctx / future serving config. The pre-P2P-patch "nan" seen at ub128 was P2P corruption, not this bug. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-DSA: convert sparse-attention control from env vars to CLI args (off by default) Implements ikawrakow's direction from discussion #2040: the DSA sparse indexer must be controllable via command-line argument (not environment variables), and must be OFF by default for now. Control surface, before -> after: DSA_INDEXER_DISABLE (env, inverted: on-by-default) -> --dsa / -dsa (cparams.dsa, default false; opt-in, dense-by-default) DSA_TOPK_OVERRIDE (env) -> --dsa-top-k N / -dsatk N (cparams.dsa_top_k, default -1 == model's configured indexer_top_k) DSA_HADAMARD_DISABLE, DSA_SINK (env) -> kept as DEBUG-ONLY env knobs (clearly commented; no CLI surface, not system on/off controls) Plumbing mirrors existing boolean/int feature flags (-mla, -khad): include/llama.h llama_context_params {bool dsa; int dsa_top_k;} src/llama.cpp default_params (false / -1); cparams assignment src/llama-cparams.h llama_cparams {bool dsa=false; int dsa_top_k=-1;} common/common.h gpt_params {bool dsa=false; int dsa_top_k=-1;} common/common.cpp arg parse + help text + cparams copy src/graphs/build_deepseek2.cpp gate now checks cparams.dsa instead of getenv; top-k override reads cparams.dsa_top_k. Stays arch-gated to LLM_ARCH_GLM_DSA. When --dsa is off (default) the indexer function is never called -> existing dense MLA path, byte-identical to no-feature. Validation (GLM-5.2-UD-IQ2_M, 3x P100, -ngl 99 --cpu-moe -mla 3 -fa 1, wikitext-2, 4 chunks @ c2560): --dsa OFF (default, dense): PPL 2.4151 (graph nodes 4166) --dsa ON, default top_k=2048: PPL 2.4697 (graph nodes 8846) --dsa ON, --dsa-top-k 1024: PPL 3.5107 Off-by-default runs the dense path; ON activates the indexer (node count jumps, PPL shifts as the top-k mask bites once n_kv > top_k). No env var is consulted for the primary on/off or the top-k knob. Graph-parallel (-sm graph) interaction (the item ikawrakow flagged): Under -sm graph the MLA layers are TP-split (wo->extra) and route to build_deepseek2_tp_attention(), which contains NO indexer code. So --dsa is silently a NO-OP under -sm graph: it does not error or crash, it runs dense. Empirically, --dsa --dsa-top-k 1024 under -sm graph gives PPL 2.4308 (chunks 1.6967/1.7906/2.1664/2.4308) -- the dense baseline (2.4151), NOT the DSA top_k=1024 numbers (3.5107). The 0.016 delta is f16 TP-reduce numerics, not DSA. Conclusion: DSA "works under deepseek2" only on the non-TP (layer) path; serving DSA with -sm graph would require wiring the indexer into the TP attention path (or a dedicated DSA arch). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-DSA: warn that --dsa is inactive under -sm graph/attn (TP path runs dense MLA) The DSA lightning indexer is built only in the layer-mode (non-TP) attention path. Under -sm graph / -sm attn the tensor-parallel attention path has no indexer, so --dsa would silently run dense MLA. Emit a clear one-time LLAMA_LOG_WARN at context creation instead of degrading silently. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-DSA: drop in-tree dev reference docs from the PR branch DSA_REFERENCE.md and the R740 progress note are development scratch, not part of the submission. Remove them so the PR diff is code-only. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * GLM-DSA: fix CPU-only crashes in the sparse-attention path PR #2045 adds GLM-DSA sparse attention but was validated on CUDA (--cpu-moe). A CPU-only build (-ngl 0 --dsa) crashes in four spots where the CUDA backend tolerates something the CPU backend does not. These make GLM-5.2 --dsa run coherently on CPU; with --dsa off they are no-ops (DSA CPU path only). 1. set_rows into an F32 dest segfaults (ggml.c set_rows_f32): type_traits[F32].from_float is NULL, so the DSA sparse-mask scatter calls a NULL fn (segfault at ip=0). memcpy when the dest is F32. CUDA has a real F32 set_rows path, so this only bit the CPU build. 2. ggml_add(F32 score, F16 mask) aborts on CPU (build_deepseek2_dsa_indexer and build_deepseek2_dsa_sparse_mask): under -fa 1 the dense KQ_mask is F16 and CPU add only accepts F32+F16 when src0 is F16. Cast the causal mask view to F32. CUDA's add accepts the mixed types. 3. dsa_fa_mask dim-1 concat must be F32 on CPU (build_deepseek2_dsa_fa_mask): CPU ggml_concat only supports F16 along dim 0; do the row (dim-1) concat in F32 then cast the result to F16. CUDA supports the F16 dim-1 concat. 4. indexer k_norm epsilon is 0 -> ggml_norm aborts (llama-hparams.cpp): the lightning-indexer k_norm is a non-RMS LayerNorm using f_norm_eps, but the GLM-DSA GGUF only carries the RMS eps so f_norm_eps stays 0 (GGML_ASSERT(eps > 0)). Mirror the RMS eps. CUDA's norm doesn't assert on eps=0. Validated: GLM-5.2 UD-Q4_K_M, single-socket Xeon w7-2475X, CPU-only (-ngl 0 --dsa) - coherent at 49K+ ctx, correct 30K needle retrieval, prefill flat with length (~32 tok/s, the O(L) DSA signature) vs the dense build's O(L^2) decline. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * DSA: loop over attention heads + use builtin Hadamard * DSA: ggml_blend * DSA: remove a bunch of unnecessary ggml_cont * DSA: fix CUDA blend - but something is still wrong * DSA: use ggml_top_k instead of ggml_argsort when FA is ON * CUDA: add CUB based argsort * DSA: avoid graph leaves * Various * GLM-5.2 DSA: IndexShare (shared layers reuse full-layer top-k) GLM-5.2's indexer_types marks 21 'full' layers that compute their own lightning-indexer top-k and 57 'shared' layers that reuse the previous full layer's top-k. This port computed an independent top-k on every layer, which mis-selects keys on the 57 shared layers (the transformers reference sets indexer=None on shared layers and reuses prev_topk). Shared layers now reuse the most-recent full layer's selection. Full/ shared map derived from the config rule (full iff il<=1 or il%4==2), which reproduces indexer_types exactly; loader can later override from GGUF metadata. Built on #2063's tree; head-loop/ggml_hadamard/ggml_blend/ argsort/FA-mask unchanged. 4K PPL (unsloth IQ2_M, top_k 2048, CPU): DSA-on 3.1922 -> 2.7111, dense 2.6972 (~97% of the gap). top_k>=n_kv reproduces dense exactly. Single- seq and 4x8 parallel decode coherent. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Apply suggestion from @ikawrakow --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com> Co-authored-by: mgkwill <168222+mgkwill@users.noreply.github.com> Co-authored-by: Kawrakow <iwankawrakow@gmail.com>
1676 lines
87 KiB
C++
1676 lines
87 KiB
C++
|
|
#include "llama-hparams.h"
|
|
#include "llama-model-loader.h"
|
|
#include "llama-model.h"
|
|
|
|
#include <limits>
|
|
#include <map>
|
|
|
|
#define LLAMA_MAX_EXPERTS 512 // Qwen3 Next
|
|
|
|
static const std::map<llama_rope_scaling_type, const char *> LLAMA_ROPE_SCALING_TYPES = {
|
|
{ LLAMA_ROPE_SCALING_TYPE_NONE, "none" },
|
|
{ LLAMA_ROPE_SCALING_TYPE_LINEAR, "linear" },
|
|
{ LLAMA_ROPE_SCALING_TYPE_YARN, "yarn" },
|
|
};
|
|
|
|
static llama_rope_scaling_type llama_rope_scaling_type_from_string(const std::string & name) {
|
|
for (const auto & kv : LLAMA_ROPE_SCALING_TYPES) {
|
|
if (kv.second == name) {
|
|
return (llama_rope_scaling_type) kv.first;
|
|
}
|
|
}
|
|
|
|
return LLAMA_ROPE_SCALING_TYPE_UNSPECIFIED;
|
|
}
|
|
|
|
const char * llama_hparams::rope_scaling_type_name(llama_rope_scaling_type type) {
|
|
return LLAMA_ROPE_SCALING_TYPES.at(type);
|
|
}
|
|
|
|
static inline const char * llm_expert_gating_func_name(llm_expert_gating_func_type type) {
|
|
switch (type) {
|
|
case LLM_EXPERT_GATING_FUNC_SOFTMAX: return "softmax";
|
|
case LLM_EXPERT_GATING_FUNC_SIGMOID: return "sigmoid";
|
|
case LLM_EXPERT_GATING_FUNC_TYPE_SOFTMAX_WEIGHT: return "weight";
|
|
default: return "none";
|
|
}
|
|
}
|
|
|
|
static bool load_dflash_target_layer_ids(
|
|
llama_model_loader & ml,
|
|
const std::string & key,
|
|
llama_hparams & hparams,
|
|
bool required) {
|
|
const int kid = gguf_find_key(ml.meta, key.c_str());
|
|
if (kid < 0 || gguf_get_kv_type(ml.meta, kid) != GGUF_TYPE_ARRAY) {
|
|
if (required) {
|
|
throw std::runtime_error(format("array key not found in model: %s", key.c_str()));
|
|
}
|
|
return false;
|
|
}
|
|
|
|
const enum gguf_type type = gguf_get_arr_type(ml.meta, kid);
|
|
if (type != GGUF_TYPE_UINT32 && type != GGUF_TYPE_INT32) {
|
|
throw std::runtime_error(format("dflash: %s must be a uint32/int32 array", key.c_str()));
|
|
}
|
|
|
|
uint32_t n = 0;
|
|
ml.get_arr_n(key, n, true);
|
|
if (n == 0) {
|
|
throw std::runtime_error(format("dflash: %s must not be empty", key.c_str()));
|
|
}
|
|
if (n > 8) {
|
|
throw std::runtime_error(format("dflash: %s has %u entries, max is 8", key.c_str(), n));
|
|
}
|
|
|
|
hparams.dflash_n_target_layers = n;
|
|
for (uint32_t & id : hparams.dflash_target_layer_ids) {
|
|
id = 0;
|
|
}
|
|
|
|
if (type == GGUF_TYPE_INT32) {
|
|
std::array<int32_t, 8> layer_ids = {};
|
|
ml.get_arr(key, layer_ids, true);
|
|
for (uint32_t i = 0; i < hparams.dflash_n_target_layers; ++i) {
|
|
if (layer_ids[i] < 0) {
|
|
throw std::runtime_error(format("dflash: %s contains negative layer id %d", key.c_str(), layer_ids[i]));
|
|
}
|
|
hparams.dflash_target_layer_ids[i] = (uint32_t) layer_ids[i];
|
|
}
|
|
} else {
|
|
std::array<uint32_t, 8> layer_ids = {};
|
|
ml.get_arr(key, layer_ids, true);
|
|
for (uint32_t i = 0; i < hparams.dflash_n_target_layers; ++i) {
|
|
hparams.dflash_target_layer_ids[i] = layer_ids[i];
|
|
}
|
|
}
|
|
|
|
for (uint32_t i = 0; i < hparams.dflash_n_target_layers; ++i) {
|
|
const uint32_t id = hparams.dflash_target_layer_ids[i];
|
|
for (uint32_t j = 0; j < i; ++j) {
|
|
if (hparams.dflash_target_layer_ids[j] == id) {
|
|
throw std::runtime_error(format(
|
|
"dflash: %s contains duplicate layer id %u",
|
|
key.c_str(),
|
|
id));
|
|
}
|
|
}
|
|
}
|
|
|
|
return true;
|
|
}
|
|
|
|
static void validate_dflash_hparams(llama_hparams & hparams, llm_arch arch) {
|
|
if (hparams.dflash_block_size <= 1) {
|
|
throw std::runtime_error(format("%s: dflash block_size must be > 1", llama_model_arch_name(arch)));
|
|
}
|
|
if (hparams.dflash_n_target_layers == 0) {
|
|
throw std::runtime_error(format("%s: dflash target_layer_ids are required", llama_model_arch_name(arch)));
|
|
}
|
|
|
|
// DFlash feature width is target-model specific. Keep the serialized metadata intact here
|
|
// and validate it against the live target model during DFlash init.
|
|
|
|
if (hparams.dflash_n_target_features == 0) {
|
|
throw std::runtime_error(format(
|
|
"%s: dflash n_target_features must be > 0",
|
|
llama_model_arch_name(arch)));
|
|
}
|
|
if (hparams.dflash_n_target_features % hparams.dflash_n_target_layers != 0) {
|
|
throw std::runtime_error(format(
|
|
"%s: dflash n_target_features=%u must be divisible by n_target_layers=%u",
|
|
llama_model_arch_name(arch),
|
|
hparams.dflash_n_target_features,
|
|
hparams.dflash_n_target_layers));
|
|
}
|
|
}
|
|
|
|
|
|
void llm_load_hparams(
|
|
llama_model_loader & ml,
|
|
llama_model & model, bool ignore_vocab) {
|
|
auto & hparams = model.hparams;
|
|
const gguf_context * ctx = ml.meta;
|
|
|
|
// get metadata as string
|
|
for (int i = 0; i < gguf_get_n_kv(ctx); i++) {
|
|
enum gguf_type type = gguf_get_kv_type(ctx, i);
|
|
if (type == GGUF_TYPE_ARRAY) {
|
|
continue;
|
|
}
|
|
const char * name = gguf_get_key(ctx, i);
|
|
const std::string value = gguf_kv_to_str(ctx, i);
|
|
model.gguf_kv.emplace(name, value);
|
|
}
|
|
|
|
ml.get_key(LLM_KV_BLOCK_COUNT, hparams.n_layer);
|
|
|
|
// get general kv
|
|
ml.get_key(LLM_KV_GENERAL_NAME, model.name, false);
|
|
|
|
// get hparams kv
|
|
ml.get_key(LLM_KV_VOCAB_SIZE, hparams.n_vocab, false) || ml.get_arr_n(LLM_KV_TOKENIZER_LIST, hparams.n_vocab, !ignore_vocab);
|
|
|
|
// everything past this point is not vocab-related
|
|
if (hparams.vocab_only) {
|
|
return;
|
|
}
|
|
|
|
ml.get_key(LLM_KV_CONTEXT_LENGTH, hparams.n_ctx_train);
|
|
ml.get_key(LLM_KV_EMBEDDING_LENGTH, hparams.n_embd);
|
|
ml.get_key(LLM_KV_EXPERT_COUNT, hparams.n_expert, false);
|
|
ml.get_key(LLM_KV_EXPERT_USED_COUNT, hparams.n_expert_used, false);
|
|
|
|
GGML_ASSERT(hparams.n_expert <= LLAMA_MAX_EXPERTS);
|
|
GGML_ASSERT(hparams.n_expert_used <= hparams.n_expert);
|
|
if (hparams.n_expert > 0) {
|
|
GGML_ASSERT(hparams.n_expert_used > 0);
|
|
} else {
|
|
GGML_ASSERT(hparams.n_expert_used == 0);
|
|
}
|
|
|
|
// zero-out the per-layer hparams
|
|
std::fill(hparams.n_head_arr.begin(), hparams.n_head_arr.end(), 0);
|
|
std::fill(hparams.n_head_kv_arr.begin(), hparams.n_head_kv_arr.end(), 0);
|
|
std::fill(hparams.n_ff_arr.begin(), hparams.n_ff_arr.end(), 0);
|
|
std::fill(hparams.swa_layers.begin(), hparams.swa_layers.end(), 0);
|
|
std::fill(hparams.recurrent_layer_arr.begin(), hparams.recurrent_layer_arr.end(), false);
|
|
|
|
ml.get_key_or_arr(LLM_KV_FEED_FORWARD_LENGTH, hparams.n_ff_arr, hparams.n_layer, false);
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_HEAD_COUNT, hparams.n_head_arr, hparams.n_layer, false);
|
|
|
|
// n_head_kv is optional, default to n_head
|
|
hparams.n_head_kv_arr = hparams.n_head_arr;
|
|
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_HEAD_COUNT_KV, hparams.n_head_kv_arr, hparams.n_layer, false);
|
|
|
|
bool rope_finetuned = false;
|
|
ml.get_key(LLM_KV_ROPE_SCALING_FINETUNED, rope_finetuned, false);
|
|
hparams.rope_finetuned = rope_finetuned;
|
|
|
|
hparams.n_ctx_orig_yarn = hparams.n_ctx_train;
|
|
ml.get_key(LLM_KV_ROPE_SCALING_ORIG_CTX_LEN, hparams.n_ctx_orig_yarn, false);
|
|
|
|
// rope_freq_base (optional)
|
|
hparams.rope_freq_base_train = 10000.0f;
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE, hparams.rope_freq_base_train, false);
|
|
|
|
std::string rope_scaling("linear");
|
|
ml.get_key(LLM_KV_ROPE_SCALING_TYPE, rope_scaling, false);
|
|
hparams.rope_scaling_type_train = llama_rope_scaling_type_from_string(rope_scaling);
|
|
GGML_ASSERT(hparams.rope_scaling_type_train != LLAMA_ROPE_SCALING_TYPE_UNSPECIFIED);
|
|
|
|
// rope_freq_scale (inverse of the kv) is optional
|
|
float ropescale = 0.0f;
|
|
if (!ml.get_key(LLM_KV_ROPE_SCALING_FACTOR, ropescale, false)) {
|
|
// try the old key name
|
|
ml.get_key(LLM_KV_ROPE_SCALE_LINEAR, ropescale, false);
|
|
}
|
|
hparams.rope_freq_scale_train = ropescale == 0.0f ? 1.0f : 1.0f/ropescale;
|
|
|
|
// by default assume that the sliding-window layers use the same scaling type as the non-sliding-window layers
|
|
hparams.rope_freq_base_train_swa = hparams.rope_freq_base_train;
|
|
hparams.rope_freq_scale_train_swa = hparams.rope_freq_scale_train;
|
|
|
|
ml.get_key(LLM_KV_ROPE_SCALING_ATTN_FACTOR, hparams.rope_attn_factor, false);
|
|
|
|
// non-transformer models do not have attention heads
|
|
if (hparams.n_head() > 0) {
|
|
// gpt-neox n_rot = rotary_pct * (n_embd / n_head)
|
|
// gpt-j n_rot = rotary_dim
|
|
|
|
hparams.n_embd_head_k_full = hparams.n_embd / hparams.n_head();
|
|
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH, hparams.n_embd_head_k_full, false);
|
|
|
|
hparams.n_embd_head_v_full = hparams.n_embd / hparams.n_head();
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH, hparams.n_embd_head_v_full, false);
|
|
|
|
// sanity check for n_rot (optional)
|
|
hparams.n_rot = hparams.n_embd_head_k_full;
|
|
|
|
ml.get_key(LLM_KV_ROPE_DIMENSION_COUNT, hparams.n_rot, false);
|
|
|
|
if (model.arch == LLM_ARCH_LLAMA || model.arch == LLM_ARCH_FALCON || model.arch == LLM_ARCH_BITNET_25 || model.arch == LLM_ARCH_BITNET_B158 || model.arch == LLM_ARCH_DECI) {
|
|
if (hparams.n_rot != hparams.n_embd_head_k_full) {
|
|
throw std::runtime_error(format("invalid n_rot: %u, expected %u", hparams.n_rot, hparams.n_embd_head_k_full));
|
|
}
|
|
}
|
|
} else {
|
|
hparams.n_rot = 0;
|
|
hparams.n_embd_head_k_full = 0;
|
|
hparams.n_embd_head_v_full = 0;
|
|
}
|
|
|
|
{
|
|
hparams.n_embd_head_k_swa = hparams.n_embd_head_k_full;
|
|
hparams.n_embd_head_v_swa = hparams.n_embd_head_v_full;
|
|
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH_SWA, hparams.n_embd_head_k_swa, false);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH_SWA, hparams.n_embd_head_v_swa, false);
|
|
|
|
hparams.n_rot_swa = hparams.n_rot;
|
|
ml.get_key(LLM_KV_ROPE_DIMENSION_COUNT_SWA, hparams.n_rot_swa, false);
|
|
}
|
|
|
|
// arch-specific KVs
|
|
switch (model.arch) {
|
|
case LLM_ARCH_LLAMA:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
if (hparams.n_expert == 8) {
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_8x7B; break;
|
|
case 56: model.type = e_model::MODEL_8x22B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} else {
|
|
switch (hparams.n_layer) {
|
|
case 22: model.type = e_model::MODEL_1B; break;
|
|
case 26: model.type = e_model::MODEL_3B; break;
|
|
// granite uses a vocab with len 49152
|
|
case 32: model.type = hparams.n_vocab == 49152 ? e_model::MODEL_3B : (hparams.n_vocab < 40000 ? e_model::MODEL_7B : e_model::MODEL_8B); break;
|
|
case 36: model.type = e_model::MODEL_8B; break; // granite
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
case 48: model.type = e_model::MODEL_34B; break;
|
|
case 60: model.type = e_model::MODEL_30B; break;
|
|
case 80: model.type = hparams.n_head() == hparams.n_head_kv() ? e_model::MODEL_65B : e_model::MODEL_70B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
}
|
|
} break;
|
|
case LLM_ARCH_LLAMA4:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_INTERLEAVE_MOE_LAYER_STEP, hparams.n_moe_layer_step);
|
|
hparams.n_swa_pattern = 4; // pattern: 3 chunked - 1 full
|
|
hparams.n_attn_chunk = 8192; // should this be a gguf kv? currently it's the same for Scout and Maverick
|
|
hparams.n_swa = 1; // TODO @ngxson : this is added to trigger the SWA branch (we store the chunked attn mask in the SWA tensor), will need to clean this up later
|
|
|
|
switch (hparams.n_expert) {
|
|
case 16: model.type = MODEL_17B_16E; break;
|
|
case 128: model.type = MODEL_17B_128E; break;
|
|
default: model.type = MODEL_UNKNOWN;
|
|
}
|
|
|
|
if (model.type == MODEL_17B_128E) {
|
|
hparams.use_kq_norm = false;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_DECI:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 80: model.type = e_model::MODEL_70B; break;
|
|
case 162: model.type = e_model::MODEL_405B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MINICPM:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_2B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GROK:
|
|
{
|
|
// defaults for old GGUFs
|
|
hparams.yarn_beta_fast = 8.0f;
|
|
hparams.f_logit_scale = 0.5773502691896257f;
|
|
hparams.f_embedding_scale = 78.38367176906169f;
|
|
hparams.f_attn_out_scale = 0.08838834764831845f;
|
|
hparams.f_attn_logit_softcapping = 30.0f;
|
|
hparams.f_router_logit_softcapping = 30.0f;
|
|
// no final_logit_softcapping in grok-1
|
|
hparams.f_final_logit_softcapping = 0.0f;
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_LOGIT_SCALE, hparams.f_logit_scale, false);
|
|
ml.get_key(LLM_KV_EMBEDDING_SCALE, hparams.f_embedding_scale, false);
|
|
ml.get_key(LLM_KV_ATTENTION_OUTPUT_SCALE, hparams.f_attn_out_scale, false);
|
|
ml.get_key(LLM_KV_ATTN_LOGIT_SOFTCAPPING, hparams.f_attn_logit_softcapping, false);
|
|
ml.get_key(LLM_KV_ROUTER_LOGIT_SOFTCAPPING, hparams.f_router_logit_softcapping, false);
|
|
ml.get_key(LLM_KV_FINAL_LOGIT_SOFTCAPPING, hparams.f_final_logit_softcapping, false);
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_TEMPERATURE_LENGTH, hparams.attn_temp_length, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_EXT_FACTOR, hparams.yarn_ext_factor, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_ATTN_FACTOR, hparams.yarn_attn_factor, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_FAST, hparams.yarn_beta_fast, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_SLOW, hparams.yarn_beta_slow, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 64: model.type = e_model::MODEL_314B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_FALCON:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 60: model.type = e_model::MODEL_40B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_BAICHUAN:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
|
|
if (model.type == e_model::MODEL_13B) {
|
|
// TODO: become GGUF KV parameter
|
|
hparams.f_max_alibi_bias = 8.0f;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_STARCODER:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1B; break;
|
|
case 36: model.type = e_model::MODEL_3B; break;
|
|
case 42: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = e_model::MODEL_15B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_REFACT:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_1B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
|
|
// TODO: become GGUF KV parameter
|
|
hparams.f_max_alibi_bias = 8.0f;
|
|
} break;
|
|
case LLM_ARCH_BERT:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CAUSAL, hparams.causal_attn);
|
|
ml.get_key(LLM_KV_TOKENIZER_TOKEN_TYPE_COUNT, hparams.n_vocab_type);
|
|
ml.get_key(LLM_KV_POOLING_TYPE, hparams.pooling_type, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 3:
|
|
model.type = e_model::MODEL_17M; break; // bge-micro
|
|
case 6:
|
|
model.type = e_model::MODEL_22M; break; // MiniLM-L6
|
|
case 12:
|
|
switch (hparams.n_embd) {
|
|
case 384: model.type = e_model::MODEL_33M; break; // MiniLM-L12, bge-small
|
|
case 768: model.type = e_model::MODEL_109M; break; // bge-base
|
|
} break;
|
|
case 24:
|
|
model.type = e_model::MODEL_335M; break; // bge-large
|
|
}
|
|
} break;
|
|
case LLM_ARCH_JINA_BERT_V2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CAUSAL, hparams.causal_attn);
|
|
ml.get_key(LLM_KV_TOKENIZER_TOKEN_TYPE_COUNT, hparams.n_vocab_type);
|
|
ml.get_key(LLM_KV_POOLING_TYPE, hparams.pooling_type);
|
|
hparams.f_max_alibi_bias = 8.0f;
|
|
|
|
switch (hparams.n_layer) {
|
|
case 4: model.type = e_model::MODEL_33M; break; // jina-embeddings-small
|
|
case 12: model.type = e_model::MODEL_137M; break; // jina-embeddings-base
|
|
}
|
|
} break;
|
|
case LLM_ARCH_NOMIC_BERT:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CAUSAL, hparams.causal_attn);
|
|
ml.get_key(LLM_KV_TOKENIZER_TOKEN_TYPE_COUNT, hparams.n_vocab_type);
|
|
ml.get_key(LLM_KV_POOLING_TYPE, hparams.pooling_type);
|
|
|
|
if (hparams.n_layer == 12 && hparams.n_embd == 768) {
|
|
model.type = e_model::MODEL_137M;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_BLOOM:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1B; break;
|
|
case 30:
|
|
switch (hparams.n_embd) {
|
|
case 2560: model.type = e_model::MODEL_3B; break;
|
|
case 4096: model.type = e_model::MODEL_7B; break;
|
|
} break;
|
|
}
|
|
|
|
// TODO: become GGUF KV parameter
|
|
hparams.f_max_alibi_bias = 8.0f;
|
|
} break;
|
|
case LLM_ARCH_MPT:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CLAMP_KQV, hparams.f_clamp_kqv, false);
|
|
ml.get_key(LLM_KV_ATTENTION_MAX_ALIBI_BIAS, hparams.f_max_alibi_bias);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 48: model.type = e_model::MODEL_30B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_STABLELM:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1B; break;
|
|
case 32: model.type = e_model::MODEL_3B; break;
|
|
case 40: model.type = e_model::MODEL_12B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN2VL:
|
|
{
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, true);
|
|
}
|
|
// fall through
|
|
case LLM_ARCH_QWEN2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = hparams.n_embd == 1024 ? e_model::MODEL_0_5B : e_model::MODEL_1B; break;
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = hparams.n_head() == 20 ? e_model::MODEL_4B : e_model::MODEL_13B; break;
|
|
case 80: model.type = e_model::MODEL_70B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN2MOE:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_A2_7B; break;
|
|
case 28: model.type = e_model::MODEL_57B_A14B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
|
|
case LLM_ARCH_QWEN3:
|
|
{
|
|
ml.get_key(LLM_KV_POOLING_TYPE, hparams.pooling_type, false);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 28: model.type = hparams.n_embd == 1024 ? e_model::MODEL_0_6B : e_model::MODEL_1_7B; break;
|
|
case 36: model.type = hparams.n_embd == 2560 ? e_model::MODEL_4B : e_model::MODEL_8B; break;
|
|
case 40: model.type = e_model::MODEL_14B; break;
|
|
case 64: model.type = e_model::MODEL_32B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN3VL:
|
|
{
|
|
ml.get_key(LLM_KV_NUM_DEEPSTACK_LAYERS, hparams.n_deepstack_layers, false);
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, true);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 28: model.type = e_model::MODEL_1_7B; break;
|
|
case 36: model.type = hparams.n_embd == 2560 ? e_model::MODEL_4B : e_model::MODEL_8B; break;
|
|
case 64: model.type = e_model::MODEL_32B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
// since vision model stacks deepstack features along feature dim
|
|
// we also create a fake "n_embd" for text model to be the main embd + deepstack embds
|
|
hparams.n_embd *= hparams.n_deepstack_layers + 1;
|
|
} break;
|
|
case LLM_ARCH_QWEN3MOE:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 48: model.type = e_model::MODEL_30B_A3B; break;
|
|
case 94: model.type = e_model::MODEL_235B_A22B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MELLUM:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa, false);
|
|
|
|
if (hparams.n_swa > 0) {
|
|
hparams.rope_freq_base_train_swa = hparams.rope_freq_base_train;
|
|
hparams.rope_freq_scale_train_swa = 1; //hparams.rope_freq_scale_train;
|
|
|
|
if (!ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer, false)) {
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
hparams.swa_layers[i] = ((i + 1) % 4 != 0);
|
|
}
|
|
}
|
|
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa, false);
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 28: model.type = e_model::MODEL_12B_A2_5B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN3NEXT:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
ml.get_key(LLM_KV_SSM_CONV_KERNEL, hparams.ssm_d_conv);
|
|
ml.get_key(LLM_KV_SSM_INNER_SIZE, hparams.ssm_d_inner);
|
|
ml.get_key(LLM_KV_SSM_STATE_SIZE, hparams.ssm_d_state);
|
|
ml.get_key(LLM_KV_SSM_TIME_STEP_RANK, hparams.ssm_dt_rank);
|
|
ml.get_key(LLM_KV_SSM_GROUP_COUNT, hparams.ssm_n_group);
|
|
|
|
// Upstream convention: every 4th layer is full attention, others are recurrent.
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
hparams.recurrent_layer_arr[i] = ((i + 1) % 4 != 0);
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 48: model.type = e_model::MODEL_80B_A3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN35MOE:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, true);
|
|
|
|
ml.get_key(LLM_KV_NEXTN_PREDICT_LAYERS, hparams.nextn_predict_layers, false);
|
|
if (model.mtp) {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
} else {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - hparams.nextn_predict_layers;
|
|
}
|
|
|
|
// Load linear attention (gated delta net) parameters
|
|
ml.get_key(LLM_KV_SSM_CONV_KERNEL, hparams.ssm_d_conv);
|
|
ml.get_key(LLM_KV_SSM_INNER_SIZE, hparams.ssm_d_inner);
|
|
ml.get_key(LLM_KV_SSM_STATE_SIZE, hparams.ssm_d_state);
|
|
ml.get_key(LLM_KV_SSM_TIME_STEP_RANK, hparams.ssm_dt_rank);
|
|
ml.get_key(LLM_KV_SSM_GROUP_COUNT, hparams.ssm_n_group);
|
|
|
|
// Mark recurrent layers (linear attention layers)
|
|
{
|
|
uint32_t full_attn_interval = 4;
|
|
ml.get_key(LLM_KV_FULL_ATTENTION_INTERVAL, full_attn_interval, false);
|
|
const uint32_t n_main_layers = hparams.n_layer - hparams.nextn_predict_layers;
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
if (i < n_main_layers) {
|
|
hparams.recurrent_layer_arr[i] = ((i + 1) % full_attn_interval != 0);
|
|
} else {
|
|
hparams.recurrent_layer_arr[i] = false;
|
|
}
|
|
}
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40:
|
|
case 41:
|
|
model.type = e_model::MODEL_35B_A3B; break;
|
|
case 48: model.type = e_model::MODEL_122B_A10B; break;
|
|
case 60: model.type = e_model::MODEL_397B_A17B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN35:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, true);
|
|
|
|
// NextN/MTP parameters
|
|
ml.get_key(LLM_KV_NEXTN_PREDICT_LAYERS, hparams.nextn_predict_layers, false);
|
|
if (model.mtp) {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
} else {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - hparams.nextn_predict_layers;
|
|
}
|
|
|
|
// Load linear attention (gated delta net) parameters
|
|
ml.get_key(LLM_KV_SSM_CONV_KERNEL, hparams.ssm_d_conv);
|
|
ml.get_key(LLM_KV_SSM_INNER_SIZE, hparams.ssm_d_inner);
|
|
ml.get_key(LLM_KV_SSM_STATE_SIZE, hparams.ssm_d_state);
|
|
ml.get_key(LLM_KV_SSM_TIME_STEP_RANK, hparams.ssm_dt_rank);
|
|
ml.get_key(LLM_KV_SSM_GROUP_COUNT, hparams.ssm_n_group);
|
|
|
|
// Mark recurrent layers (linear attention layers)
|
|
// MTP layers always use standard attention, not delta-net
|
|
{
|
|
uint32_t full_attn_interval = 4;
|
|
ml.get_key(LLM_KV_FULL_ATTENTION_INTERVAL, full_attn_interval, false);
|
|
const uint32_t n_main_layers = hparams.n_layer - hparams.nextn_predict_layers;
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
if (i < n_main_layers) {
|
|
hparams.recurrent_layer_arr[i] = ((i + 1) % full_attn_interval != 0);
|
|
} else {
|
|
hparams.recurrent_layer_arr[i] = false;
|
|
}
|
|
}
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: // without MTP layer
|
|
case 25: // with MTP layer (24 main + 1 MTP)
|
|
model.type = hparams.n_embd == 1024 ? e_model::MODEL_0_8B : e_model::MODEL_2B; break;
|
|
case 32: // without MTP layer
|
|
case 33: // with MTP layer (32 main + 1 MTP)
|
|
model.type = hparams.n_embd == 2560 ? e_model::MODEL_4B : e_model::MODEL_9B; break;
|
|
case 64: // without MTP layer
|
|
case 65: // with MTP layer (64 main + 1 MTP)
|
|
model.type = e_model::MODEL_27B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_QWEN3VLMOE:
|
|
{
|
|
ml.get_key(LLM_KV_NUM_DEEPSTACK_LAYERS, hparams.n_deepstack_layers, false);
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, true);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 48: model.type = e_model::MODEL_30B_A3B; break;
|
|
case 94: model.type = e_model::MODEL_235B_A22B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
// since vision model stacks deepstack features along feature dim
|
|
// we also create a fake "n_embd" for text model to be the main embd + deepstack embds
|
|
hparams.n_embd *= hparams.n_deepstack_layers + 1;
|
|
} break;
|
|
case LLM_ARCH_PHI2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1B; break;
|
|
case 32: model.type = e_model::MODEL_3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_PHI3:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1B; break;
|
|
case 32: model.type = e_model::MODEL_3B; break;
|
|
case 40: model.type = e_model::MODEL_14B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
|
|
// for backward compatibility ; see: https://github.com/ggerganov/llama.cpp/pull/8931
|
|
if ((hparams.n_layer == 32 || hparams.n_layer == 40) && hparams.n_ctx_train == 4096) {
|
|
// default value for Phi-3-mini-4k-instruct and Phi-3-medium-4k-instruct
|
|
hparams.n_swa = 2047;
|
|
} else if (hparams.n_layer == 32 && hparams.n_head_kv(0) == 32 && hparams.n_ctx_train == 131072) {
|
|
// default value for Phi-3-mini-128k-instruct
|
|
hparams.n_swa = 262144;
|
|
} else if (hparams.n_layer == 40 && hparams.n_ctx_train == 131072) {
|
|
// default value for Phi-3-medium-128k-instruct
|
|
hparams.n_swa = 131072;
|
|
}
|
|
bool found_swa = ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa, false);
|
|
if (!found_swa && hparams.n_swa == 0) {
|
|
throw std::runtime_error("invalid value for sliding_window");
|
|
}
|
|
} break;
|
|
case LLM_ARCH_PLAMO:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GPT2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 12: model.type = e_model::MODEL_SMALL; break;
|
|
case 24: model.type = e_model::MODEL_MEDIUM; break;
|
|
case 36: model.type = e_model::MODEL_LARGE; break;
|
|
case 48: model.type = e_model::MODEL_XL; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_CODESHELL:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 42: model.type = e_model::MODEL_7B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_ORION:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_14B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_INTERNLM2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 48: model.type = e_model::MODEL_20B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GEMMA:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 18: model.type = e_model::MODEL_2B; break;
|
|
case 28: model.type = e_model::MODEL_7B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GEMMA2:
|
|
{
|
|
hparams.n_swa = 4096; // default value of gemma 2
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa, false);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_ATTN_LOGIT_SOFTCAPPING, hparams.f_attn_logit_softcapping, false);
|
|
ml.get_key(LLM_KV_FINAL_LOGIT_SOFTCAPPING, hparams.f_final_logit_softcapping, false);
|
|
hparams.attn_soft_cap = true;
|
|
|
|
switch (hparams.n_layer) {
|
|
case 26: model.type = e_model::MODEL_2B; break;
|
|
case 42: model.type = e_model::MODEL_9B; break;
|
|
case 46: model.type = e_model::MODEL_27B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GEMMA3:
|
|
{
|
|
hparams.n_swa_pattern = 6;
|
|
|
|
hparams.rope_freq_base_train_swa = 10000.0f;
|
|
hparams.rope_freq_scale_train_swa = 1.0f;
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 26: model.type = e_model::MODEL_2B; break;
|
|
case 34: model.type = e_model::MODEL_4B; break;
|
|
case 48: model.type = e_model::MODEL_12B; break;
|
|
case 62: model.type = e_model::MODEL_27B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
|
|
hparams.f_attention_scale = model.type == e_model::MODEL_27B
|
|
? 1.0f / std::sqrt(float(hparams.n_embd / hparams.n_head(0)))
|
|
: 1.0f / std::sqrt(float(hparams.n_embd_head_k_full));
|
|
} break;
|
|
case LLM_ARCH_GEMMA4:
|
|
{
|
|
//hparams.swa_type = LLAMA_SWA_TYPE_STANDARD;
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer);
|
|
|
|
uint32_t n_kv_shared_layers = 0;
|
|
ml.get_key(LLM_KV_ATTENTION_SHARED_KV_LAYERS, n_kv_shared_layers, false);
|
|
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - (int32_t)n_kv_shared_layers;
|
|
hparams.f_attention_scale = 1.0f; // Gemma4 uses self.scaling = 1.0 (no pre-attn scaling)
|
|
hparams.f_final_logit_softcapping = 30.0f;
|
|
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa, false);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp, false);
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EMBEDDING_LENGTH_PER_LAYER, hparams.n_embd_per_layer);
|
|
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH_SWA, hparams.n_embd_head_k_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH_SWA, hparams.n_embd_head_v_swa);
|
|
ml.get_key(LLM_KV_FINAL_LOGIT_SOFTCAPPING, hparams.f_final_logit_softcapping, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 35: model.type = e_model::MODEL_2B; break;
|
|
case 42: model.type = e_model::MODEL_4B; break; // to confirm: E4B or E5B?
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GEMMA4_MTP:
|
|
case LLM_ARCH_GEMMA4_ASSISTANT:
|
|
{
|
|
if (model.arch == LLM_ARCH_GEMMA4_MTP) {
|
|
ml.get_key(LLM_KV_MTP_BACKBONE_EMBEDDING_LENGTH, hparams.mtp_backbone_n_embd);
|
|
ml.get_key(LLM_KV_MTP_CENTROID_COUNT, hparams.mtp_num_centroids, false);
|
|
ml.get_key(LLM_KV_MTP_CENTROID_TOP_K, hparams.mtp_centroid_top_k, false);
|
|
} else {
|
|
ml.get_key("gemma4-assistant.embedding_length_out", hparams.mtp_backbone_n_embd);
|
|
ml.get_key("gemma4-assistant.n_centroids", hparams.mtp_num_centroids, false);
|
|
ml.get_key("gemma4-assistant.centroid_top_k", hparams.mtp_centroid_top_k, false);
|
|
}
|
|
ml.get_key(LLM_KV_MTP_USE_ORDERED_EMBEDDINGS, hparams.mtp_use_ordered_embeddings, false);
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer);
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa, false);
|
|
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
hparams.f_attention_scale = 1.0f;
|
|
|
|
switch (hparams.mtp_backbone_n_embd) {
|
|
case 5376: model.type = e_model::MODEL_32B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_DFLASH_DRAFT:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_DFLASH_BLOCK_SIZE, hparams.dflash_block_size, false);
|
|
ml.get_key(LLM_KV_DFLASH_MASK_TOKEN_ID, hparams.dflash_mask_token_id, false);
|
|
ml.get_key(LLM_KV_DFLASH_N_TARGET_FEATURES, hparams.dflash_n_target_features, false);
|
|
ml.get_key(LLM_KV_DFLASH_BACKBONE_ROTARY_BASE, hparams.dflash_backbone_rotary_base, false);
|
|
load_dflash_target_layer_ids(ml, LLM_KV(model.arch)(LLM_KV_DFLASH_TARGET_LAYER_IDS), hparams, false);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_SCALE, hparams.f_attn_v_scale, false);
|
|
// DFlash drafts may be trained with sliding-window attention (for long-context).
|
|
// Read the window + per-layer pattern so the SWA mask path activates; absent keys
|
|
// leave n_swa=0 / swa_layers all-zero (dense behavior, unchanged).
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa, false);
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer, false);
|
|
validate_dflash_hparams(hparams, model.arch);
|
|
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
|
|
case LLM_ARCH_STARCODER2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 30: model.type = e_model::MODEL_3B; break;
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = e_model::MODEL_15B; break;
|
|
case 52: model.type = e_model::MODEL_20B; break; // granite
|
|
case 88: model.type = e_model::MODEL_34B; break; // granite
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MAMBA:
|
|
{
|
|
ml.get_key(LLM_KV_SSM_CONV_KERNEL, hparams.ssm_d_conv);
|
|
ml.get_key(LLM_KV_SSM_INNER_SIZE, hparams.ssm_d_inner);
|
|
ml.get_key(LLM_KV_SSM_STATE_SIZE, hparams.ssm_d_state);
|
|
ml.get_key(LLM_KV_SSM_TIME_STEP_RANK, hparams.ssm_dt_rank);
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24:
|
|
switch (hparams.n_embd) {
|
|
case 768: model.type = e_model::MODEL_SMALL; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 48:
|
|
switch (hparams.n_embd) {
|
|
case 1024: model.type = e_model::MODEL_MEDIUM; break;
|
|
case 1536: model.type = e_model::MODEL_LARGE; break;
|
|
case 2048: model.type = e_model::MODEL_XL; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 64:
|
|
switch (hparams.n_embd) {
|
|
case 2560: model.type = e_model::MODEL_3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_XVERSE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
case 80: model.type = e_model::MODEL_65B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_COMMAND_R:
|
|
{
|
|
ml.get_key(LLM_KV_LOGIT_SCALE, hparams.f_logit_scale);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_35B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_DBRX:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CLAMP_KQV, hparams.f_clamp_kqv);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_16x12B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_OLMO:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_CLAMP_KQV, hparams.f_clamp_kqv, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 22: model.type = e_model::MODEL_1B; break;
|
|
case 32: model.type = e_model::MODEL_7B; break;
|
|
case 80: model.type = e_model::MODEL_70B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_OPENELM:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 16: model.type = e_model::MODEL_270M; break;
|
|
case 20: model.type = e_model::MODEL_450M; break;
|
|
case 28: model.type = e_model::MODEL_1B; break;
|
|
case 36: model.type = e_model::MODEL_3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GPTNEOX:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_USE_PARALLEL_RESIDUAL, hparams.use_par_res);
|
|
switch (hparams.n_layer) {
|
|
case 6:
|
|
switch (hparams.n_ff()) {
|
|
case 512: model.type = e_model::MODEL_14M; break;
|
|
case 2048: model.type = e_model::MODEL_70M; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 12:
|
|
switch (hparams.n_ff()) {
|
|
case 3072: model.type = e_model::MODEL_160M; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 16:
|
|
switch (hparams.n_ff()) {
|
|
case 8192: model.type = e_model::MODEL_1B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 24:
|
|
switch (hparams.n_ff()) {
|
|
case 4096: model.type = e_model::MODEL_410M; break;
|
|
case 8192: model.type = e_model::MODEL_1_4B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 32:
|
|
switch (hparams.n_ff()) {
|
|
case 10240: model.type = e_model::MODEL_2_8B; break;
|
|
case 16384: model.type = e_model::MODEL_6_9B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 36:
|
|
switch (hparams.n_ff()) {
|
|
case 20480: model.type = e_model::MODEL_12B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 44:
|
|
switch (hparams.n_ff()) {
|
|
case 24576: model.type = e_model::MODEL_20B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_ARCTIC:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
if (hparams.n_expert == 128) {
|
|
switch (hparams.n_layer) {
|
|
case 35: model.type = e_model::MODEL_10B_128x3_66B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} else {
|
|
model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MISTRAL4:
|
|
case LLM_ARCH_DEEPSEEK2:
|
|
{
|
|
int expected_head_size_k = model.arch == LLM_ARCH_DEEPSEEK2 ? 576 : 320;
|
|
int expected_head_size_v = model.arch == LLM_ARCH_DEEPSEEK2 ? 512 : 256;
|
|
if (hparams.n_head_kv() == 1) {
|
|
int n_nead_kv = hparams.n_gqa();
|
|
if (n_nead_kv%4 != 0 || hparams.n_embd_head_k(0) != expected_head_size_k || hparams.n_embd_head_v(0) != expected_head_size_v ||
|
|
hparams.n_rot != 64) {
|
|
LLAMA_LOG_ERROR("==========================================================================\n");
|
|
LLAMA_LOG_ERROR("Detected incompatible DeepSeek model without a known way to fix it.\n");
|
|
LLAMA_LOG_ERROR("Consider making your own ik_llama.cpp compatible model or\n");
|
|
LLAMA_LOG_ERROR("ask the model provider to make one for you,\n\n");
|
|
LLAMA_LOG_ERROR("Sorry, uknown model => cannot fix it => bailing out\n");
|
|
LLAMA_LOG_ERROR("==========================================================================\n");
|
|
GGML_ABORT("Fatal error");
|
|
}
|
|
LLAMA_LOG_INFO("================= Adjusted mainline llama.cpp MLA tensors to ik_llama.cpp\n");
|
|
for (auto& item : hparams.n_head_kv_arr) item = n_nead_kv;
|
|
hparams.n_embd_head_k_full = 192;
|
|
hparams.n_embd_head_v_full = 128;
|
|
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH_MLA, hparams.n_embd_head_k_full);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH_MLA, hparams.n_embd_head_v_full);
|
|
}
|
|
bool is_lite = (hparams.n_layer == 27 || hparams.n_layer == 26) || (hparams.n_layer == 48 && hparams.n_vocab == 128256);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead);
|
|
if (!is_lite) {
|
|
ml.get_key(LLM_KV_ATTENTION_Q_LORA_RANK, hparams.n_lora_q);
|
|
}
|
|
ml.get_key(LLM_KV_ATTENTION_KV_LORA_RANK, hparams.n_lora_kv);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_TYPE_NONE;
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
if (hparams.expert_gating_func == LLM_EXPERT_GATING_FUNC_TYPE_NONE) {
|
|
// Older DeepSeek models from the 2.0/2.5 series may not have the experts gating function recorded in the GGUF.
|
|
// Such models use SOFTMAX as the experts gating function.
|
|
// The new (new as of this commit) GLM-4.7-Flash may also be missing the experts gating function.
|
|
// GLM-4.7-Flash uses SIGMOID as the experts gating function.
|
|
// Hence, we make the LLM_KV_EXPERT_GATING_FUNC entry optional, and set here if missing.
|
|
// We distinguish between GLM-4.7-Flash and DeepSeek-2/2.5 models by the number of layers.
|
|
// GLM-4.7-Flash has 47 layers (or 48, if an MTP layer is included in the GGUF).
|
|
hparams.expert_gating_func = hparams.n_layer == 47 || hparams.n_layer == 48 ?
|
|
LLM_EXPERT_GATING_FUNC_SIGMOID : LLM_EXPERT_GATING_FUNC_SOFTMAX;
|
|
LLAMA_LOG_INFO("================= Missing experts gating function -> set to %s\n",
|
|
llm_expert_gating_func_name(llm_expert_gating_func_type(hparams.expert_gating_func)));
|
|
}
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_LOG_MUL, hparams.rope_yarn_log_mul, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 27: model.type = e_model::MODEL_16B; break;
|
|
case 36: model.type = e_model::MODEL_119B_A6B; break;
|
|
case 47: model.type = e_model::MODEL_30B_A3B; break; // GLM-4.7-Flash
|
|
case 60: model.type = e_model::MODEL_236B; break;
|
|
case 61: model.type = e_model::MODEL_671B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_CHATGLM:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 28: model.type = e_model::MODEL_6B; break;
|
|
case 40: model.type = e_model::MODEL_9B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GLM4:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_9B; break;
|
|
case 61: model.type = e_model::MODEL_32B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GLM4_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
// MoE parameters
|
|
ml.get_key(LLM_KV_EXPERT_COUNT, hparams.n_expert);
|
|
ml.get_key(LLM_KV_EXPERT_USED_COUNT, hparams.n_expert_used);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
|
|
// Expert gating function (GLM4_MOE uses sigmoid)
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
if (hparams.expert_gating_func == 0) {
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_SIGMOID;
|
|
}
|
|
|
|
// NextN/MTP parameters
|
|
if (model.mtp) {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
}
|
|
else {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - hparams.nextn_predict_layers;
|
|
}
|
|
ml.get_key(LLM_KV_NEXTN_PREDICT_LAYERS, hparams.nextn_predict_layers, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 47: model.type = e_model::MODEL_106B_A12B; break; // GLM-4.5-Air (46 layers + 1 NextN layer)
|
|
case 93: model.type = e_model::MODEL_355B_A32B; break; // GLM-4.5 (92 layers + 1 NextN layer)
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_BITNET:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 26: model.type = e_model::MODEL_3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_BITNET_B158:
|
|
case LLM_ARCH_BITNET_25:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 30: model.type = e_model::MODEL_2B; break; // bitnet2b_2501
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_T5:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_RELATIVE_BUCKETS_COUNT, hparams.n_rel_attn_bkts);
|
|
|
|
uint32_t dec_start_token_id;
|
|
if (ml.get_key(LLM_KV_DECODER_START_TOKEN_ID, dec_start_token_id, false)) {
|
|
hparams.dec_start_token_id = dec_start_token_id;
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 6: model.type = e_model::MODEL_60M; break; // t5-small
|
|
case 8: model.type = e_model::MODEL_80M; break; // flan-t5-small
|
|
case 12:
|
|
switch (hparams.n_ff()) {
|
|
case 3072: model.type = e_model::MODEL_220M; break; // t5-base
|
|
case 2048: model.type = e_model::MODEL_250M; break; // flan-t5-base
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case 24:
|
|
switch (hparams.n_ff()) {
|
|
case 4096: model.type = e_model::MODEL_770M; break; // t5-large
|
|
case 2816: model.type = e_model::MODEL_780M; break; // flan-t5-large
|
|
case 16384: model.type = e_model::MODEL_3B; break; // t5-3b
|
|
case 5120: model.type = e_model::MODEL_3B; break; // flan-t5-xl
|
|
case 65536: model.type = e_model::MODEL_11B; break; // t5-11b
|
|
case 10240: model.type = e_model::MODEL_11B; break; // flan-t5-xxl
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_T5ENCODER:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_RELATIVE_BUCKETS_COUNT, hparams.n_rel_attn_bkts);
|
|
model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case LLM_ARCH_JAIS:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_MAX_ALIBI_BIAS, hparams.f_max_alibi_bias);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 24: model.type = e_model::MODEL_1_3B; break;
|
|
case 40: model.type = e_model::MODEL_13B; break;
|
|
/* TODO: add variants */
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GRANITE:
|
|
case LLM_ARCH_GRANITE_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LOGIT_SCALE, hparams.f_logit_scale);
|
|
ml.get_key(LLM_KV_RESIDUAL_SCALE, hparams.f_residual_scale);
|
|
ml.get_key(LLM_KV_EMBEDDING_SCALE, hparams.f_embedding_scale);
|
|
ml.get_key(LLM_KV_ATTENTION_SCALE, hparams.f_attention_scale);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_3B; break;
|
|
case 40: model.type = e_model::MODEL_3B; break;
|
|
// Add additional layer/vocab/etc checks here for other model sizes
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_COHERE2:
|
|
{
|
|
hparams.n_swa_pattern = 4;
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_LOGIT_SCALE, hparams.f_logit_scale);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_EPS, hparams.f_norm_eps);
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_8B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_COHERE2_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer);
|
|
ml.get_key(LLM_KV_LOGIT_SCALE, hparams.f_logit_scale);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
if (hparams.expert_gating_func == LLM_EXPERT_GATING_FUNC_TYPE_NONE) {
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_SIGMOID;
|
|
}
|
|
switch (hparams.n_layer) {
|
|
case 49: model.type = e_model::MODEL_30B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_BAILINGMOE2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_EXPERT_GROUP_COUNT, hparams.n_expert_groups);
|
|
ml.get_key(LLM_KV_EXPERT_GROUP_USED_COUNT, hparams.n_group_used);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func);
|
|
ml.get_key(LLM_KV_NEXTN_PREDICT_LAYERS, hparams.nextn_predict_layers, false);
|
|
|
|
// TODO: when MTP is implemented, this should probably be updated if needed
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - hparams.nextn_predict_layers;
|
|
|
|
switch (hparams.n_layer) {
|
|
case 20: model.type = MODEL_16B_A1B; break;
|
|
case 21: model.type = MODEL_16B_A1B; break;
|
|
case 32: model.type = MODEL_100B_A6B; break;
|
|
case 33: model.type = MODEL_100B_A6B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_DOTS1:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
switch (hparams.n_layer) {
|
|
case 62: model.type = e_model::MODEL_142B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_ERNIE4_5:
|
|
case LLM_ARCH_ERNIE4_5_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
if (model.arch == LLM_ARCH_ERNIE4_5_MOE) {
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
ml.get_key(LLM_KV_INTERLEAVE_MOE_LAYER_STEP, hparams.n_moe_layer_step);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead);
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 18: model.type = e_model::MODEL_0_3B; break;
|
|
case 28: model.type = e_model::MODEL_21B_A3B; break;
|
|
case 54: model.type = e_model::MODEL_300B_A47B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_HUNYUAN_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 32: model.type = e_model::MODEL_80B_A13B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_OPENAI_MOE:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
|
|
//TODO OAI_MOE: SWA
|
|
//hparams.swa_type = LLAMA_SWA_TYPE_STANDARD;
|
|
//hparams.set_swa_pattern(2);
|
|
|
|
// TODO: switch (hparams.n_layer)
|
|
|
|
} break;
|
|
case LLM_ARCH_MINIMAX_M2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 62: model.type = e_model::MODEL_230B_A10B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MINIMAX_M3:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead, false);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
|
|
model.type = e_model::MODEL_UNKNOWN;
|
|
} break;
|
|
case LLM_ARCH_SMOLLM3:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
hparams.n_no_rope_layer_step = 4;
|
|
|
|
switch (hparams.n_layer) {
|
|
case 36: model.type = e_model::MODEL_3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MISTRAL3:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
ml.get_key(LLM_KV_ATTENTION_TEMPERATURE_SCALE, hparams.f_attn_temp_scale, false);
|
|
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_FAST, hparams.yarn_beta_fast, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_SLOW, hparams.yarn_beta_slow, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_LOG_MUL, hparams.rope_yarn_log_mul, false);
|
|
|
|
if (hparams.f_attn_temp_scale != 0.0f) {
|
|
hparams.n_attn_temp_floor_scale = hparams.n_ctx_orig_yarn;
|
|
if (hparams.n_attn_temp_floor_scale == 0) {
|
|
throw std::runtime_error("invalid n_ctx_orig_yarn for attention temperature scaling");
|
|
}
|
|
}
|
|
|
|
// TODO: this seems to be correct with the case of mscale == mscale_all_dims == 1.0f
|
|
// but may need further verification with other values
|
|
if (hparams.rope_yarn_log_mul != 0.0f) {
|
|
float factor = 1.0f / hparams.rope_freq_scale_train;
|
|
float mscale = 1.0f;
|
|
float mscale_all_dims = hparams.rope_yarn_log_mul;
|
|
static auto get_mscale = [](float scale, float mscale) {
|
|
return scale <= 1.0f ? 1.0f : (0.1f * mscale * logf(scale) + 1.0f);
|
|
};
|
|
hparams.yarn_attn_factor = get_mscale(factor, mscale) / get_mscale(factor, mscale_all_dims);
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 26: model.type = e_model::MODEL_3B; break;
|
|
case 34: model.type = e_model::MODEL_8B; break;
|
|
case 40: model.type = e_model::MODEL_14B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_MIMO2:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_SCALE, hparams.f_attn_v_scale, false);
|
|
//TODO
|
|
//hparams.swa_type = LLAMA_SWA_TYPE_STANDARD; // which is the same as OpenAI
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer);
|
|
|
|
switch (hparams.n_layer) {
|
|
case 48: model.type = e_model::MODEL_310B_A15B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
|
|
} break;
|
|
case LLM_ARCH_SEED_OSS:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
switch (hparams.n_layer) {
|
|
case 64: model.type = e_model::MODEL_36B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_STEP35:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
//hparams.swa_type = LLAMA_SWA_TYPE_STANDARD;
|
|
// MoE + SWA parameters
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
// Step35 uses sigmoid gating by default (if not set in GGUF)
|
|
if (hparams.expert_gating_func == LLM_EXPERT_GATING_FUNC_TYPE_NONE) {
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_SIGMOID;
|
|
}
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
bool have_rfb_train_swa = ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa, false);
|
|
ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer);
|
|
if (!ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_COUNT_PER_LAYER, hparams.rope_dim_per_layer, hparams.n_layer, false)) {
|
|
for (int i = 0; i < hparams.n_layer; ++i) {
|
|
hparams.rope_dim_per_layer[i] = hparams.swa_layers[i] ? hparams.n_rot : hparams.n_rot/2;
|
|
}
|
|
}
|
|
// The following two parameters: one of the two versions must be present in the GGUF
|
|
if (!ml.get_key_or_arr(LLM_KV_SWIGLU_LIMITS, hparams.swiglu_limits, hparams.n_layer, false)) {
|
|
ml.get_key_or_arr(LLM_KV_SWIGLU_CLAMP_EXP, hparams.swiglu_limits, hparams.n_layer, true);
|
|
}
|
|
if (!ml.get_key_or_arr(LLM_KV_SWIGLU_LIMITS_SHARED, hparams.swiglu_limits_shared, hparams.n_layer, false)) {
|
|
ml.get_key_or_arr(LLM_KV_SWIGLU_CLAMP_SHEXP, hparams.swiglu_limits_shared, hparams.n_layer, true);
|
|
}
|
|
// Optional: Step35-only gating for applying rope scaling (HF: yarn_only_types).
|
|
// Default is 3 (apply on all layers) if the key is absent.
|
|
//ml.get_key(format("%s.rope.scaling.apply_mask", ml.get_arch_name().c_str()),
|
|
// hparams.rope_scaling_apply_mask, false);
|
|
//hparams.has_rope_freq_base_per_layer = ml.get_key_or_arr(
|
|
// format("%s.rope.freq_base_per_layer", ml.get_arch_name().c_str()),
|
|
// hparams.rope_freq_base_per_layer, hparams.n_layer, false);
|
|
ml.get_key(format("%s.rope.scaling.apply_mask", ml.get_arch_name().c_str()),
|
|
hparams.rope_scaling_apply_mask, false);
|
|
hparams.has_rope_freq_base_per_layer = ml.get_key_or_arr(LLM_KV_ROPE_FREQ_BASE_PER_LAYER,
|
|
hparams.rope_freq_base_per_layer, hparams.n_layer, false);
|
|
GGML_ASSERT(hparams.has_rope_freq_base_per_layer || have_rfb_train_swa);
|
|
} break;
|
|
case LLM_ARCH_LAGUNA:
|
|
{
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_FEED_FORWARD_LENGTH, hparams.n_ff_shexp, false);
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_TYPE_NONE;
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead, false);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared, false);
|
|
// Older Laguna GGUFs encode one shared expert through the shared FFN length.
|
|
if (hparams.n_expert_shared == 0 && hparams.n_ff_shexp > 0) {
|
|
hparams.n_expert_shared = 1;
|
|
}
|
|
if (hparams.expert_gating_func == LLM_EXPERT_GATING_FUNC_TYPE_NONE) {
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_SIGMOID;
|
|
}
|
|
|
|
ml.get_key(LLM_KV_ATTENTION_SLIDING_WINDOW, hparams.n_swa);
|
|
ml.get_key(LLM_KV_ROPE_FREQ_BASE_SWA, hparams.rope_freq_base_train_swa, false);
|
|
hparams.rope_freq_scale_train_swa = 1.0f;
|
|
if (!ml.get_key_or_arr(LLM_KV_ATTENTION_SLIDING_WINDOW_PATTERN, hparams.swa_layers, hparams.n_layer, false)) {
|
|
// Laguna XS.2 alternates full-attention and SWA layers via per-layer head counts.
|
|
const uint32_t n_head_full = hparams.n_head(0);
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
hparams.swa_layers[i] = hparams.n_head(i) != n_head_full;
|
|
}
|
|
}
|
|
|
|
const bool found_rope_dim = ml.get_key(LLM_KV_ROPE_DIMENSION_COUNT, hparams.n_rot, false);
|
|
const bool found_rope_dim_swa = ml.get_key(LLM_KV_ROPE_DIMENSION_COUNT_SWA, hparams.n_rot_swa, false);
|
|
|
|
// Laguna GGUFs store the number of scalar Q/K dimensions that ggml_rope_ext
|
|
// rotates. Correct files carry those values explicitly. Some early public
|
|
// XS.2 GGUFs omitted both keys, so fall back to the HF XS.2 layout only for
|
|
// missing metadata: full-attention layers rotate half the head, SWA layers
|
|
// rotate the full head. Explicit but wrong halved metadata still needs repair.
|
|
if (hparams.n_swa > 0) {
|
|
if (!found_rope_dim) {
|
|
hparams.n_rot = hparams.n_embd_head_k_full / 2;
|
|
}
|
|
if (!found_rope_dim_swa) {
|
|
hparams.n_rot_swa = hparams.n_embd_head_k_swa;
|
|
}
|
|
}
|
|
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_EXT_FACTOR, hparams.yarn_ext_factor, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_ATTN_FACTOR, hparams.yarn_attn_factor, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_FAST, hparams.yarn_beta_fast, false);
|
|
ml.get_key(LLM_KV_ROPE_SCALING_YARN_BETA_SLOW, hparams.yarn_beta_slow, false);
|
|
|
|
if (!ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_COUNT_PER_LAYER, hparams.rope_dim_per_layer, hparams.n_layer, false)) {
|
|
for (uint32_t i = 0; i < hparams.n_layer; ++i) {
|
|
hparams.rope_dim_per_layer[i] = hparams.swa_layers[i] ? hparams.n_rot_swa : hparams.n_rot;
|
|
}
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 40: model.type = e_model::MODEL_33B_A3B; break;
|
|
default: model.type = e_model::MODEL_UNKNOWN;
|
|
}
|
|
} break;
|
|
case LLM_ARCH_GLM_DSA:
|
|
{
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_ATTENTION_LAYERNORM_RMS_EPS, hparams.f_norm_rms_eps);
|
|
// GLM-DSA lightning-indexer k_norm is a (non-RMS) LayerNorm built via LLM_NORM,
|
|
// which uses hparams.f_norm_eps in ggml_norm(). The GGUF only carries the RMS eps,
|
|
// so f_norm_eps stays 0 and CPU ggml_norm aborts (GGML_ASSERT(eps > 0)). On CUDA
|
|
// the kernel does not assert (eps=0 is numerically tolerable), which is why the
|
|
// CPU attention path was never exercised. Mirror the RMS eps so the indexer
|
|
// LayerNorm gets a valid epsilon on all backends. (HF hardcodes 1e-6 for this
|
|
// LayerNorm, but 1e-6 vs 1e-5 is within 4-chunk PPL noise here, so keep the mirror.)
|
|
if (hparams.f_norm_eps <= 0.0f) hparams.f_norm_eps = hparams.f_norm_rms_eps;
|
|
ml.get_key_or_arr(LLM_KV_ROPE_DIMENSION_SECTIONS, hparams.rope_sections, 4, false);
|
|
|
|
// MoE parameters
|
|
ml.get_key(LLM_KV_EXPERT_COUNT, hparams.n_expert);
|
|
ml.get_key(LLM_KV_EXPERT_USED_COUNT, hparams.n_expert_used);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
ml.get_key(LLM_KV_LEADING_DENSE_BLOCK_COUNT, hparams.n_layer_dense_lead, false);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_SCALE, hparams.expert_weights_scale);
|
|
ml.get_key(LLM_KV_EXPERT_WEIGHTS_NORM, hparams.expert_weights_norm, false);
|
|
|
|
// deepseek MLA parameters
|
|
ml.get_key(LLM_KV_ATTENTION_Q_LORA_RANK, hparams.n_lora_q);
|
|
ml.get_key(LLM_KV_ATTENTION_KV_LORA_RANK, hparams.n_lora_kv);
|
|
//ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH_MLA, hparams.n_embd_head_k_mla_impl, false);
|
|
//ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH_MLA, hparams.n_embd_head_v_mla_impl, false);
|
|
ml.get_key(LLM_KV_EXPERT_FEED_FORWARD_LENGTH, hparams.n_ff_exp);
|
|
ml.get_key(LLM_KV_EXPERT_SHARED_COUNT, hparams.n_expert_shared);
|
|
|
|
// DSA parameters
|
|
ml.get_key(LLM_KV_ATTENTION_INDEXER_HEAD_COUNT, hparams.indexer_n_head);
|
|
ml.get_key(LLM_KV_ATTENTION_INDEXER_KEY_LENGTH, hparams.indexer_head_size);
|
|
ml.get_key(LLM_KV_ATTENTION_INDEXER_TOP_K, hparams.indexer_top_k);
|
|
|
|
// GLM-5.2 IndexShare: per-layer full/shared indexer map. "full" layers compute their own
|
|
// top-k; "shared" layers reuse the previous full layer's selection (transformers
|
|
// modeling_glm_moe_dsa.py: shared layer indexer=None, topk_indices=prev_topk_indices).
|
|
// Derived from GLM-5.2's config indexer_types rule (full iff il<=1 or il%4==2), verified
|
|
// to reproduce the config's full set {0,1,2,6,10,...} exactly. Existing GGUFs carry no
|
|
// per-layer metadata, so the derivation is the source of truth; a future metadata key can
|
|
// override this for GLM-DSA variants with a different pattern.
|
|
for (uint32_t il = 0; il < hparams.n_layer; ++il) {
|
|
hparams.indexer_is_full[il] = (il <= 1) || (il % 4 == 2);
|
|
}
|
|
|
|
// Expert gating function (GLM-4.5 uses sigmoid)
|
|
ml.get_key(LLM_KV_EXPERT_GATING_FUNC, hparams.expert_gating_func, false);
|
|
if (hparams.expert_gating_func == LLM_EXPERT_GATING_FUNC_TYPE_NONE) {
|
|
hparams.expert_gating_func = LLM_EXPERT_GATING_FUNC_SIGMOID;
|
|
}
|
|
|
|
// NextN/MTP parameters
|
|
ml.get_key(LLM_KV_NEXTN_PREDICT_LAYERS, hparams.nextn_predict_layers, false);
|
|
|
|
if (model.mtp) {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer;
|
|
}
|
|
else {
|
|
hparams.n_layer_kv_from_start = hparams.n_layer - hparams.nextn_predict_layers;
|
|
}
|
|
|
|
switch (hparams.n_layer) {
|
|
case 79: model.type = MODEL_744B_A40B; break;
|
|
default: model.type = MODEL_UNKNOWN;
|
|
}
|
|
if (hparams.n_head_kv() == 1) {
|
|
int n_nead_kv = hparams.n_gqa();
|
|
if (n_nead_kv%4 != 0 || hparams.n_embd_head_k_full != 576 || hparams.n_embd_head_v_full != 512 ||
|
|
hparams.n_rot != 64) {
|
|
LLAMA_LOG_ERROR("==========================================================================\n");
|
|
LLAMA_LOG_ERROR("Detected incompatible DeepSeek model without a known way to fix it.\n");
|
|
LLAMA_LOG_ERROR("Sorry, uknown model => cannot fix it => bailing out\n");
|
|
LLAMA_LOG_ERROR("==========================================================================\n");
|
|
GGML_ABORT("Fatal error");
|
|
}
|
|
LLAMA_LOG_INFO("================= Adjusted mainline llama.cpp MLA tensors to ik_llama.cpp\n");
|
|
for (auto& item : hparams.n_head_kv_arr) item = n_nead_kv;
|
|
hparams.n_embd_head_k_full = 192;
|
|
hparams.n_embd_head_v_full = 128;
|
|
ml.get_key(LLM_KV_ATTENTION_KEY_LENGTH_MLA, hparams.n_embd_head_k_full);
|
|
ml.get_key(LLM_KV_ATTENTION_VALUE_LENGTH_MLA, hparams.n_embd_head_v_full);
|
|
}
|
|
} break;
|
|
default: (void)0;
|
|
}
|
|
|
|
model.ftype = ml.ftype;
|
|
|
|
if (hparams.f_max_alibi_bias > 0.0f) {
|
|
hparams.use_alibi = true;
|
|
}
|
|
|
|
hparams.rope_type = llama_rope_type(&model);
|
|
}
|