Default Branch

9d07d8681e · perplexity: fix int overflows in large context x vocab buffer sizing (#2150) · Updated 2026-07-18 06:13:47 +00:00

Branches

c24d50dd88 · Split mode graph for MiniMax-M3 · Updated 2026-06-15 08:41:34 +00:00

140
0
Included

c73bfbe9ce · Fix #1961 · Updated 2026-06-14 07:42:39 +00:00

146
0
Included

175819b4fb · Style · Updated 2026-06-12 06:19:06 +00:00

154
0
Included

c622ea37d3 · More info · Updated 2026-06-11 14:06:07 +00:00

157
2

0b3d85fe3a · Minor · Updated 2026-06-10 14:48:15 +00:00

159
2

decfaf4dd3 · A few more named nodes · Updated 2026-06-10 09:31:31 +00:00

161
2

2741330db5 · Adjust CUDA FA kernel parameters for head size 512 on Turing · Updated 2026-06-09 15:04:10 +00:00

165
1

14e960489c · More · Updated 2026-06-09 13:05:24 +00:00

165
2

44e8cf1ab3 · Split mode graph for Laguna · Updated 2026-06-09 05:33:55 +00:00

168
1

17f05fc6ec · Fix bf16 graph reduce type · Updated 2026-06-08 13:35:20 +00:00

170
1

1a685f1af1 · Support for alternative Gemma4 assistant · Updated 2026-06-08 12:17:26 +00:00

170
1

98cf5cef72 · CPU FA: disable mask optimization · Updated 2026-06-08 06:18:41 +00:00

170
1

9b4b9ca4ae · CUDA FA: cover Gemma4-4B/2B assistant · Updated 2026-06-08 05:37:15 +00:00

173
1

c3b975eb04 · CPU FA: Check for empty attention mask · Updated 2026-06-05 09:21:15 +00:00

175
1

68a94ab930 · Enable split mode graph for Gemma4-12B · Updated 2026-06-04 16:12:53 +00:00

177
1

0ad43359a4 · Split mode graph for Mellum · Updated 2026-06-04 13:13:33 +00:00

182
1

adeff7dbd3 · Add extra nodes when dealing with MLA and amb · Updated 2026-05-29 08:38:59 +00:00

186
1

ccc48d33c7 · quantize: add exception for Gemma4 · Updated 2026-05-29 06:10:15 +00:00

187
1

20c2f6d97f · Add lower bound to the -amb command line argument · Updated 2026-05-29 05:22:20 +00:00

188
1

6648aa2e6e · Fix Gemma4 vision · Updated 2026-05-28 15:08:46 +00:00

187
0
Included