| Title: | Vendored 'GGML' Tensor Library with C-Callable Compute API |
|---|---|
| Description: | Vendors the core, official 'GGUF' implementation, and CPU backend of the 'GGML' tensor library (<https://github.com/ggml-org/ggml>) as a static library and exposes its core tensor-context and matrix-multiply compute path through 'R_RegisterCCallable' C-callable entry points. This is a carrier package: it has no high-level R modeling API of its own. Other R packages link to it ('LinkingTo') to build and compute 'GGML' tensor graphs (including quantized types such as Q4_K) from their own C or C++ code without re-vendoring 'GGML'. The CPU backend is always built. Native targets also build the BLAS backend for dense products; webR uses the CPU backend because its Fortran ABI differs. The 'Vulkan' and NVIDIA 'CUDA' backends are independent opt-in builds. |
| Authors: | Sounkou Mahamane Toure [aut, cre], Georgi Gerganov [cph] (Author of the GGML library), The ggml.ai / llama.cpp contributors [cph] (GGML CPU backend contributors), Yuri Baramykov [ctb] (Author of the ggmlR R/CRAN-compliance I/O shim adapted here) |
| Maintainer: | Sounkou Mahamane Toure <[email protected]> |
| License: | GPL (>= 2) |
| Version: | 0.1.0 |
| Built: | 2026-07-21 19:31:58 UTC |
| Source: | https://github.com/RGenomicsETL/Rfmalloc |
Returns the version string reported by the vendored 'GGML' library at
runtime (ggml_version()), resolved through Rggml's own registered
C-callable rather than a compile-time constant, so it always reflects
the code that was actually built.
ggml_version()ggml_version()
A length-1 character vector.
ggml_version()ggml_version()
Reports the build decisions configure made, read off the preprocessor
symbols compiled into the shared object. This is deliberately not a
re-detection of the host: it answers "what is in this binary", not "what
could this machine run".
rggml_cpu_info()rggml_cpu_info()
It exists because the failure mode it guards against is invisible. A build
that misses its intended branch - falling back to GGML's portable scalar
kernels where the native NEON ones were meant, or to the dgemm_ promotion
where sgemm_ was available - compiles cleanly and passes every numerical
test, because both branches are correct. Only the speed differs. Asserting on
this list in the test suite turns those silent fallbacks into failures.
A list with
arch_kernels"arm" for GGML's hand-tuned NEON kernels,
"wasm" for its SIMD128 kernels, and "generic" for the portable
reference kernels.
simd_dispatchTRUE when the runtime CPUID dispatcher is active
(x86: the staged AVX2 q4_K variant). Always FALSE alongside
arch_kernels = "arm" or "wasm", which supersede it.
blasTRUE when GGML's BLAS backend and R's Fortran bridge
are part of this target build. It is FALSE on wasm, where webR's
hidden character-length ABI is incompatible with the native bridge.
sgemmTRUE when R's BLAS exports sgemm_ and GGML's BLAS
backend calls it directly; FALSE when the shim promotes to dgemm_
or when blas is FALSE.
vulkanTRUE when built with --with-vulkan. Whether a device
is visible is a separate question: see rggml_vulkan_info().
cudaTRUE when built with --with-cuda. Whether a device
is visible is a separate question: see rggml_cuda_info().
rggml_vulkan_info(), rggml_cuda_info()
rggml_cpu_info()rggml_cpu_info()
Reports the NVIDIA CUDA devices GGML can see. The backend is opt-in at installation because it requires the CUDA toolkit and compiles native GPU kernels:
rggml_cuda_info() rggml_has_cuda()rggml_cuda_info() rggml_has_cuda()
install.packages("Rggml", configure.args = "--with-cuda")
R CMD INSTALL --configure-args=--with-cuda .
Set CUDA_HOME when the toolkit is outside the search path. By default the
source installation targets the GPU visible at build time. Set
RGGML_CUDA_ARCH to an explicit nvcc architecture such as sm_89 or
sm_120a when building for another machine.
A build without CUDA returns zero devices rather than failing, so callers can probe and fall back to Vulkan, BLAS or CPU computation.
A list with n_devices (integer) and device (the description of
device 0, or NA when there is none).
rggml_has_cuda() returns TRUE when at least one CUDA device is
usable.
rggml_cuda_info()rggml_cuda_info()
Computes crossprod(A, B) (that is t(A) %*% B) through
'GGML' on the requested backend, using the backend-agnostic residency path:
the operands and result are allocated in one of the backend's own buffers,
the inputs uploaded, the product computed, and the result downloaded. For a
GPU backend the tensors live in device memory, so this is the entry point a
caller uses to run a dense single-precision GEMM on the GPU.
rggml_mul_mat(A, B, backend = c("cpu", "blas", "vulkan", "cuda"))rggml_mul_mat(A, B, backend = c("cpu", "blas", "vulkan", "cuda"))
A, B
|
Numeric matrices with the same number of rows (the contracted dimension). Computed in single precision on every backend. |
backend |
One of |
A numeric matrix, dim c(ncol(A), ncol(B)), equal to
crossprod(A, B) up to single-precision rounding.
rggml_vulkan_info(), rggml_cuda_info(), rggml_cpu_info()
A <- matrix(rnorm(12), 4, 3) B <- matrix(rnorm(8), 4, 2) rggml_mul_mat(A, B) # == crossprod(A, B) to fp32A <- matrix(rnorm(12), 4, 3) B <- matrix(rnorm(8), 4, 2) rggml_mul_mat(A, B) # == crossprod(A, B) to fp32
Reports the Vulkan devices GGML can see. Rggml only contains the Vulkan backend when it was built with it - it is opt-in, because generating and compiling GGML's 156 embedded SPIR-V shaders is expensive (the largest one needs several GB of RAM):
rggml_vulkan_info() rggml_has_vulkan()rggml_vulkan_info() rggml_has_vulkan()
install.packages("Rggml", configure.args = "--with-vulkan")
R CMD INSTALL --configure-args=--with-vulkan .
It also requires glslc and the Vulkan headers at build time
(libvulkan-dev + glslc on Debian/Ubuntu, or the LunarG Vulkan SDK with
VULKAN_SDK set), and a Vulkan driver at run time. A software driver such
as Mesa's lavapipe counts: it is slow, but it makes the backend testable
without a GPU.
When Rggml was built without Vulkan, this returns zero devices rather than failing, so callers can probe and fall back.
GGML_VK_ALLOW_CPU=1 opts in to CPU-type Vulkan devices such as Mesa
lavapipe. It is off by default, preserving upstream GGML's device selection.
A list with n_devices (integer) and device (the description of
device 0, or NA when there is none).
rggml_has_vulkan() returns TRUE when at least one Vulkan device
is usable.
rggml_vulkan_info()rggml_vulkan_info()