Skip to content

Pull requests: ggml-org/llama.cpp

Author
Filter by author
Loading
Label
Filter by label
Loading
Use alt + click/return to exclude labels
or + click/return for logical OR
Projects
Filter by project
Loading
Milestones
Filter by milestone
Loading
Reviews
Assignee
Filter by who’s assigned
Assigned to nobody Loading
Sort

Pull requests list

docs: Adapt conda-forge package name documentation Improvements or additions to documentation
#26229 opened Jul 28, 2026 by jjerphan Contributor Loading…
mimo2: add MTP draft support conversion model Model specific
#26228 opened Jul 28, 2026 by tnhnyzc Contributor Loading…
core : support output vocab size distinct from embedding vocab size conversion
#26226 opened Jul 28, 2026 by adithyab94 Contributor Loading…
2 tasks done
Proper fix for host buffer sync ggml changes relating to the ggml tensor library for machine learning
#26225 opened Jul 28, 2026 by pwilkin Member Loading…
metal: fix NaN in mul_mm_id when activations exceed f16 range Apple Metal https://en.wikipedia.org/wiki/Metal_(API) ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26223 opened Jul 28, 2026 by mdegans Contributor Loading…
server: abstract llama_memory calls to common_memory server
#26221 opened Jul 28, 2026 by ngxson Collaborator Loading…
vulkan: fix Raspberry Pi V3D WG≤256 / low-SMEM enablement ggml changes relating to the ggml tensor library for machine learning testing Everything test related Vulkan Issues specific to the Vulkan backend
#26214 opened Jul 28, 2026 by odilitime Loading…
server : add /slots endpoint action=clone_to (KV clone between slots) documentation Improvements or additions to documentation server
#26204 opened Jul 27, 2026 by solethais Loading…
HIP: MMQ Dispatch config modification - separation of RDNA3, 3.5 from 4 and tune 4. CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26199 opened Jul 27, 2026 by Geramy Loading…
server: fix prompt cache entry selection and f_keep filter documentation Improvements or additions to documentation server
#26198 opened Jul 27, 2026 by q-g-j Loading…
opencl: skip the Adreno KQ/KQV image kernels for multi-stream batches when FA is off ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend
#26189 opened Jul 27, 2026 by wanghqc Contributor Loading…
ggml-cpu: add -mavxvnni for clang-cl when GGML_AVX_VNNI is enabled ggml changes relating to the ggml tensor library for machine learning
#26187 opened Jul 27, 2026 by MaxCrazy1101 Loading…
model: add Kimi-K3 text model conversion model Model specific testing Everything test related
#26185 opened Jul 27, 2026 by pwilkin Member Loading…
Support quantized kv cache for Minimax M3 model Model specific
#26180 opened Jul 27, 2026 by timkhronos Contributor Loading…
Add more benchmarks to llama-eval documentation Improvements or additions to documentation examples
#26174 opened Jul 27, 2026 by pwilkin Member Loading…
tests: add model resolution test on synthetic repo listings testing Everything test related
#26172 opened Jul 27, 2026 by ServeurpersoCom Contributor Loading…
ggml-cuda: Allow transpose-free gemmv computation CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning testing Everything test related
#26171 opened Jul 27, 2026 by roberteg16 Loading…
ggml: add a scheduler sanitizer ggml changes relating to the ggml tensor library for machine learning
#26167 opened Jul 27, 2026 by am17an Contributor Loading…
[Model/VLA] Support MiniCPM-RobotManip build Compilation issues documentation Improvements or additions to documentation mtmd Related to multimodal functionality (video/image/audio) server
#26164 opened Jul 27, 2026 by tc-mb Contributor Draft
opencl: bugfix: missing increment of ref_count in ggml_backend_opencl_init() ggml changes relating to the ggml tensor library for machine learning OpenCL Issues specific to the OpenCL backend
#26162 opened Jul 27, 2026 by akleine Contributor Loading…
cuda: compact Blackwell NVFP4 MoE work scheduling (+10% to +15% prefill, NVFP4-only) CUDA Related to the CUDA backend ggml changes relating to the ggml tensor library for machine learning
#26159 opened Jul 27, 2026 by coder543 Contributor Draft
ggml: add support for MXFP8 CPU conversion ggml changes relating to the ggml tensor library for machine learning
#26157 opened Jul 27, 2026 by michaelw9999 Contributor Loading…
mtmd: support multi-row batching for deepseek-ocr mtmd Related to multimodal functionality (video/image/audio)
#26154 opened Jul 26, 2026 by ngxson Collaborator Loading…
ProTip! What’s not been updated in a month: updated:<2026-06-28.