Thor Lin
coolthor
AI & ML interests
On-device & edge LLM inference on NVIDIA GB10 / DGX Spark.
Quantization (NVFP4 W4A4/W4A16, FP8), vLLM serving, speculative
decoding (EAGLE-3), and multimodal/omni models. Measure-first
benchmarking — I publish the numbers, including the ones that fail.
Recent Activity
liked a model about 5 hours ago
Qwen/Qwen3.8-Flash-Next updated a model 22 days ago
coolthor/comfyui-zimage-sulphur-nvfp4 updated a model 22 days ago
coolthor/MiniMax-H3-pruned-NVFP4Organizations
None yet