Hugging Face
Models
Datasets
Spaces
Buckets
new
Docs
Enterprise
Pricing
Website
Tasks
HuggingChat
Collections
Languages
Organizations
Community
Blog
Posts
Daily Papers
Hardware
Learn
Discord
Forum
GitHub
Solutions
Team & Enterprise
Hugging Face PRO
Enterprise Support
Inference Providers
Inference Endpoints
Storage Buckets
Log In
Sign Up
marcsun13
/
ggml-quantization
like
0
GGUF
kernel
quantization
License:
mit
Model card
Files
Files and versions
xet
Community
Copy to bucket
new
main
ggml-quantization
319 MB
Ctrl+K
Ctrl+K
2 contributors
History:
22 commits
marcsun13
HF Staff
trim the metal shader to the dispatched kernels
87cb5c1
verified
1 day ago
build
trim the metal shader to the dispatched kernels
1 day ago
gguf_cuda
Metal: cover every quant type upstream's metallib has a gemv for
14 days ago
gguf_metal
trim the metal shader to the dispatched kernels
1 day ago
tests
Fail when a shipped artifact is not a loadable library
14 days ago
torch-ext
Metal: cover every quant type upstream's metallib has a gemv for
14 days ago
vendor
Add Metal backend; vendor whole llama.cpp trees
15 days ago
.gitattributes
4.29 kB
trim the metal shader to the dispatched kernels
1 day ago
.gitignore
Safe
95 Bytes
Minimal README: ops, devices, source and how to refresh it
15 days ago
README.md
Safe
1.69 kB
Minimal README: ops, devices, source and how to refresh it
15 days ago
SKILL.md
Safe
11.8 kB
SKILL: the LFS pointer trap, and correct the dynamo advice
14 days ago
build.toml
3.62 kB
trim the metal shader to the dispatched kernels
1 day ago
flake.lock
Safe
3.05 kB
Add Metal backend; vendor whole llama.cpp trees
15 days ago
flake.nix
Safe
290 Bytes
GGUF kernels: dequantize + fused gemv over packed blocks
17 days ago
trim_shader.py
9.1 kB
trim the metal shader to the dispatched kernels
1 day ago
vendor.py
Safe
3.29 kB
Add Metal backend; vendor whole llama.cpp trees
15 days ago