Default Branch

9d07d8681e · perplexity: fix int overflows in large context x vocab buffer sizing (#2150) · Updated 2026-07-18 06:13:47 +00:00

Branches

a18eeb01cb · Qwen3.5 MTP: extract selected tokens earlier · Updated 2026-05-28 11:50:51 +00:00

188
1

7cf668f797 · Make MTP work with split mode graph · Updated 2026-05-28 04:12:34 +00:00

189
47

5a10d701f9 · Arghh · Updated 2026-05-27 13:49:58 +00:00

189
4

68d818269e · Fix GLM MTP with split mode graph · Updated 2026-05-26 16:30:12 +00:00

193
2

5425749950 · Fix crash with GLM and MTP · Updated 2026-05-26 14:02:41 +00:00

193
1

e5732606c5 · Fix cache loading/saving for MLA models and split mode graph · Updated 2026-05-26 12:37:12 +00:00

193
1

f3e929c25e · Disable K Hadamard transform if K-head size is not a power of 2 · Updated 2026-05-26 07:19:08 +00:00

193
1

c7211cc500 · Minor logging cleanup · Updated 2026-05-23 16:05:54 +00:00

198
1

e5abe3a86a · Per GPU fit margin · Updated 2026-05-23 15:24:19 +00:00

198
1

d065b9f742 · It is actually not related to split mode graph · Updated 2026-05-23 10:47:57 +00:00

199
2

516dfb39f3 · Fix split mode graph with ngl < n_layer · Updated 2026-05-23 06:43:52 +00:00

200
1

8bf4e6ca50 · MTP tweaks 3 · Updated 2026-05-22 09:37:36 +00:00

202
1

fa1f302d77 · Fix split mode graph for Qwen35-MoE + MTP · Updated 2026-05-22 06:11:14 +00:00

203
1

0bcfde9518 · Disable split mode graph for Qwen35-MoE when MTP is enabled · Updated 2026-05-21 13:24:37 +00:00

206
1

dd123f9f4f · Fix crash with split mode graph and partial offload · Updated 2026-05-21 10:31:51 +00:00

208
1

f1e146859b · Fix Gemma4-E4B compute graph · Updated 2026-05-21 08:19:30 +00:00

208
1

a3d46a963a · Fix MTP when -no-gr is used · Updated 2026-05-20 10:36:48 +00:00

214
1

c2c12c987d · Fix mla = 1 / mla = 3 confusion · Updated 2026-05-20 07:02:00 +00:00

219
2

8b9db5efcb · Remove Makefile · Updated 2026-05-20 06:12:59 +00:00    furyhawk

216
1

301f8d9afd · Fix #1837 · Updated 2026-05-19 14:54:13 +00:00    furyhawk

220
1