forked from ggml-org/llama.cpp
-
Notifications
You must be signed in to change notification settings - Fork 69
Pull requests: thecodacus/llama.cpp
Author
Label
Projects
Milestones
Reviews
Assignee
Sort
Pull requests list
moe-cache: pick a GPU device explicitly instead of the first non-meta device
#11
opened Aug 31, 2026 by
ALeXssNdR
Loading…
docs: qwen4exp results, an upstream-sync warning, and a caveat on "bit-identical"
documentation
Improvements or additions to documentation
#10
opened Aug 27, 2026 by
cpuchip
Loading…
fix: report fit math when expert cache pack allocation fails
#7
opened Jul 25, 2026 by
thecodacus
Owner
Loading…
ProTip!
Updated in the last three days: updated:>2026-09-08.