Expert-pruned GLM-5.3 in INT4 W4A16 (compressed-tensors) for 4x H200 serving. Built on cyankiwi/GLM-5.3-AWQ-INT4. Massmax criterion.
0xSero
0xSero
AI & ML interests
Quantizing, benchmarking, training, and building.
Recent Activity
new activity 5 days ago
deepseek-ai/DeepSeek-V4.1-Flash:Running on 4x RTX PRO 6000 with NVMe offload for ngram updated a model 6 days ago
0xSero/GLM-5.3-488B-EXL3-TR3-3.42bpw published a model 6 days ago
0xSero/GLM-5.3-488B-EXL3-TR3-3.42bpwOrganizations
0xSero Releases — sorted by date
Inkling-Small EXL3 Quantization Suite
Public EXL3 trellis quantizations of thinkingmachines/Inkling-Small from 2.0 through 5.0 bpw.
-
0xSero/Inkling-Small-EXL3-2.0bpw
Image-Text-to-Text • 41B • Updated • 353 -
0xSero/Inkling-Small-EXL3-2.5bpw
Image-Text-to-Text • 49B • Updated • 16 • 1 -
0xSero/Inkling-Small-EXL3-3.0bpw
Image-Text-to-Text • 57B • Updated • 20 -
0xSero/Inkling-Small-EXL3-3.5bpw
Image-Text-to-Text • 65B • Updated • 12
Gemma — REAP
REAP-pruned Gemma-4 MoE.
Proven REAPs
Benchmarked REAP checkpoints with >=500 all-time downloads. GLM/Qwen/MiniMax/DeepSeek/Kimi/gemma.
MiniMax — REAP
REAP-pruned & quantized MiniMax-M2.1 / M2.7.
Trinity — REAP
REAP-pruned & quantized Trinity-Large-Thinking.
Datasets — observations & calibration
REAP layerwise observations, calibration sets, and training data.
GLM-5.3 REAP EXL3 Quantization Suite
Expert-pruned GLM-5.3 (753B) in EXL3. Massmax criterion, routed experts quantized, sensitive layers BF16. Ladder 500B-661B.
-
0xSero/GLM-5.3-EXL3-3.0bpw
Text Generation • 146B • Updated • 418 • 1 -
0xSero/GLM-5.3-500B-EXL3-3.0bpw
Text Generation • 99B • Updated • 590 • 16 -
0xSero/GLM-5.3-533B-EXL3-3.0bpw
Text Generation • 105B • Updated • 257 -
0xSero/GLM-5.3-569B-EXL3-3.0bpw
Text Generation • 112B • Updated • 490 • 1
Local AI Registry
Hardware, models, launch recipes, prices, raw speed sweeps, and comparable local AI benchmarks.
Qwen — REAP
REAP-pruned & quantized Qwen3.5 / 3.6 / Coder variants.
GLM — REAP
REAP-pruned & quantized GLM-4.x / 5 / 5.1 (+ Flash fine-tunes).
DeepSeek — REAP
REAP-pruned & quantized DeepSeek-V4-Flash / V3.2.
-
0xSero/DeepSeek-V3.2-345B-W3A16
Text Generation • 2B • Updated • 59 • 10 -
0xSero/DeepSeek-V3.2-508B-NVFP4
Text Generation • 2B • Updated • 27 -
0xSero/DeepSeek-V4-Flash-162B
Text Generation • 92B • Updated • 187 • 33 -
0xSero/DeepSeek-V4-Flash-162B-GGUF
Text Generation • 163B • Updated • 1.7k • 31
Nemotron — REAP
REAP-pruned & quantized NVIDIA Nemotron-3-Super.
-
0xSero/Nemotron-3-Super-64B
Text Generation • 64B • Updated • 61 • 7 -
0xSero/Nemotron-3-Super-64B-W4A16
Text Generation • 6B • Updated • 34 • 3 -
0xSero/Nemotron-3-Super-92B
Text Generation • 92B • Updated • 59 • 2 -
0xSero/Nemotron-3-Super-92B-W4A16
Text Generation • 6B • Updated • 43 • 2
Other models
Hy3, Kimi-K2.5, INTELLECT-3, NousCoder.
GLM-5.3 REAP W4A16 (Hopper)
Expert-pruned GLM-5.3 in INT4 W4A16 (compressed-tensors) for 4x H200 serving. Built on cyankiwi/GLM-5.3-AWQ-INT4. Massmax criterion.
GLM-5.3 REAP EXL3 Quantization Suite
Expert-pruned GLM-5.3 (753B) in EXL3. Massmax criterion, routed experts quantized, sensitive layers BF16. Ladder 500B-661B.
-
0xSero/GLM-5.3-EXL3-3.0bpw
Text Generation • 146B • Updated • 418 • 1 -
0xSero/GLM-5.3-500B-EXL3-3.0bpw
Text Generation • 99B • Updated • 590 • 16 -
0xSero/GLM-5.3-533B-EXL3-3.0bpw
Text Generation • 105B • Updated • 257 -
0xSero/GLM-5.3-569B-EXL3-3.0bpw
Text Generation • 112B • Updated • 490 • 1
0xSero Releases — sorted by date
Local AI Registry
Hardware, models, launch recipes, prices, raw speed sweeps, and comparable local AI benchmarks.
Inkling-Small EXL3 Quantization Suite
Public EXL3 trellis quantizations of thinkingmachines/Inkling-Small from 2.0 through 5.0 bpw.
-
0xSero/Inkling-Small-EXL3-2.0bpw
Image-Text-to-Text • 41B • Updated • 353 -
0xSero/Inkling-Small-EXL3-2.5bpw
Image-Text-to-Text • 49B • Updated • 16 • 1 -
0xSero/Inkling-Small-EXL3-3.0bpw
Image-Text-to-Text • 57B • Updated • 20 -
0xSero/Inkling-Small-EXL3-3.5bpw
Image-Text-to-Text • 65B • Updated • 12
Qwen — REAP
REAP-pruned & quantized Qwen3.5 / 3.6 / Coder variants.
Gemma — REAP
REAP-pruned Gemma-4 MoE.
GLM — REAP
REAP-pruned & quantized GLM-4.x / 5 / 5.1 (+ Flash fine-tunes).
Proven REAPs
Benchmarked REAP checkpoints with >=500 all-time downloads. GLM/Qwen/MiniMax/DeepSeek/Kimi/gemma.
DeepSeek — REAP
REAP-pruned & quantized DeepSeek-V4-Flash / V3.2.
-
0xSero/DeepSeek-V3.2-345B-W3A16
Text Generation • 2B • Updated • 59 • 10 -
0xSero/DeepSeek-V3.2-508B-NVFP4
Text Generation • 2B • Updated • 27 -
0xSero/DeepSeek-V4-Flash-162B
Text Generation • 92B • Updated • 187 • 33 -
0xSero/DeepSeek-V4-Flash-162B-GGUF
Text Generation • 163B • Updated • 1.7k • 31
MiniMax — REAP
REAP-pruned & quantized MiniMax-M2.1 / M2.7.
Nemotron — REAP
REAP-pruned & quantized NVIDIA Nemotron-3-Super.
-
0xSero/Nemotron-3-Super-64B
Text Generation • 64B • Updated • 61 • 7 -
0xSero/Nemotron-3-Super-64B-W4A16
Text Generation • 6B • Updated • 34 • 3 -
0xSero/Nemotron-3-Super-92B
Text Generation • 92B • Updated • 59 • 2 -
0xSero/Nemotron-3-Super-92B-W4A16
Text Generation • 6B • Updated • 43 • 2
Trinity — REAP
REAP-pruned & quantized Trinity-Large-Thinking.
Other models
Hy3, Kimi-K2.5, INTELLECT-3, NousCoder.
Datasets — observations & calibration
REAP layerwise observations, calibration sets, and training data.