Instructions to use mlx-community/clef-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use mlx-community/clef-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir clef-4bit mlx-community/clef-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
mlx-community/clef-4bit
Cloudflare/clef converted to MLX (4-bit) for Apple Silicon.
Clef turns a state (text, JSON, images, or video) plus a schema of typed questions into a
probability for every allowed option, in a single forward pass. It is not a chat model —
mlx_vlm.generate, mlx_lm.generate, and LM Studio will load the backbone but produce
meaningless text. Use the bundled clef_mlx.py loader, which runs the backbone and the
joint schema head.
Which variant fits your Mac?
| Variant | Base | Download | Peak memory (1k / 4k / 16k tokens) | Latency (1k / 16k tokens) | Minimum Mac RAM |
|---|---|---|---|---|---|
| clef-flash-4bit | 9B | 6.2 GB | 7.0 / 7.2 / 8.6 GB | 0.31 s / 7.0 s | 16 GB |
| clef-flash-8bit | 9B | 10.7 GB | 11.4 / 11.6 / 13.0 GB | 0.34 s / 7.7 s | 24 GB (16 GB for short prompts) |
| clef-4bit (this repo) | 27B | 16.3 GB | 17.1 / 17.5 / 19.6 GB | 1.4 s / 26.0 s | 32 GB |
| clef-8bit | 27B | 29.8 GB | 30.5 / 30.8 / 33.0 GB | 1.5 s / 32.0 s | 48 GB |
Measured on an M5 Max with text input. macOS lets the GPU use only about 70–75% of RAM by default, so the minimum RAM is higher than the peak. The 9B variants are much faster; the 27B variants score higher (see the quality check below).
Usage
pip install "mlx-vlm>=0.7.4,<0.8" huggingface_hub # no torch needed
Tested with mlx 0.32.3, mlx-lm 0.32.0 and mlx-vlm 0.7.4. clef_mlx.py uses mlx-vlm internals, so it
warns if you load it with an untested mlx-vlm minor version.
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("mlx-community/clef-4bit")
sys.path.insert(0, path)
import clef_mlx
model = clef_mlx.load(path)
response = model.systemone({
"model": "clef",
"state": "Our checkout started returning errors and orders are blocked.",
"questions": {
"department": {
"type": "choice",
"instructions": "Which team should handle the message?",
"criteria": {"billing": "Payments or invoices", "technical": "Bugs or outages"},
},
"urgency": {"type": "score", "criteria": ["Can wait", "This week", "Today"]},
"outage": {"type": "noul", "instructions": "Is a service down?"},
},
})
print(response["answers"])
Images (PIL) and videos (frame arrays) go in images / videos, as in the original:
from PIL import Image
model.predict({
"state": {"task": "Review the attached receipt."},
"images": [Image.open("receipt.jpg")],
"questions": {"legible": {"type": "noul", "instructions": "Is the receipt total legible?"}},
})
See the original model card for the input format, question types, and benchmarks.
Try it from the command line
clef_mlx.py runs on its own. From the downloaded repo it uses that repo by default; elsewhere pass
--model mlx-community/clef-4bit.
cd "$(hf download mlx-community/clef-4bit --quiet)"
python clef_mlx.py predict \
--state "Checkout has been failing for every customer for the last hour." \
--questions '{"urgent": {"type": "noul", "instructions": "Is this urgent?"},
"team": {"type": "choice", "criteria": {"billing": "Payments", "technical": "Outages and errors"}}}'
--questions takes JSON, a .json file, or - for stdin; --request takes a whole request body instead;
--image photo.jpg attaches an image (repeatable). The output is the SystemOne response as JSON.
Run a local SystemOne server
python clef_mlx.py serve --port 8000 # http://127.0.0.1:8000, local only by default
curl http://127.0.0.1:8000/v1/systemone -H "Content-Type: application/json" -d '{
"model": "clef",
"state": "Checkout has been failing for every customer for the last hour.",
"questions": {"urgent": {"type": "noul", "instructions": "Is this urgent?"}}
}'
POST /v1/systemonetakes the same request body as the Jev/SystemOne API and returns the same response (plususage.latency_ms), so existing SystemOne clients can point at your laptop.GET /healthandGET /v1/modelsare also available.- Images go in
imagesas data URLs, base64, or http(s) URLs. Videos aren't supported over HTTP. - Errors return JSON:
400for invalid requests,413with "maximum context length" for inputs that don't fit. Add"truncate": falseto a request (or start with--no-truncate) to get the413instead of truncation. - Requests run one at a time on the GPU. There is no authentication: keep the default
127.0.0.1binding unless you put your own proxy in front. - For the Decision Index, start with
--no-truncateand use itshttpengine:python -m decision_index run --engine http --option base_url=http://127.0.0.1:8000.
Limitations
- Not a chat model. Text generation tools (
mlx_lm.generate,mlx_vlm.generate, LM Studio, Ollama) load the backbone but give meaningless output. Onlyclef_mlx.pyruns the decision head. - Long inputs are truncated by default, exactly like the reference implementation: the state is cut so
the whole prompt fits in 16,384 tokens. Pass
truncate=Falsetopredict()/systemone()to get aclef_mlx.ContextTooLongerror instead, or raisemax_length(the head was trained at 16k). - One record per call. There is no batching; requests run one after another.
- Images or videos, not both in one record. Several images, or several videos, are fine.
- Speed depends heavily on the chip. Latencies above are from an M5 Max; long prompts on base M-series chips will be several times slower.
- Custom code. Like the original release, this repo ships Python (
clef_mlx.py) that you import and run. Read it before use if that matters in your environment.
Conversion
- Backbone:
mlx_vlm.convert -q --q-bits 4 --q-group-size 64(vision tower kept in bf16). - Joint schema head:
joint_head.safetensorscopied unchanged (bf16) and run byclef_mlx.py. processor_config.jsonis the original from Cloudflare/clef; prompt/token layout matches the referencejoint_schema_model.pyexactly (images and video).
Parity vs. official PyTorch implementation (bf16)
| Inputs | Top answer agrees | Max abs Δprob |
|---|---|---|
| Text (4 records, 10 questions) | 10/10 | 0.037 |
| Images + video (5 records, 9 questions) | 9/9 | 0.097 |
Measured on an M5 Max (128 GB). Small spot-check, not a full benchmark run.
Quality check: Decision Index (sampled)
| Decision Index (sample) | Same top answer as 8-bit | Mean max abs Δp | Median latency | |
|---|---|---|---|---|
| This model (4-bit) | 57.02 | 98.1% (15,915 answers) | 0.023 | 1.03 s |
| MLX 8-bit, same rows | 57.96 | — | — | 1.13 s |
| Cloudflare published (full suite) | 61.21 |
4-bit costs about 1 index point vs 8-bit on identical rows (8-bit tracks the PyTorch reference within 0.007 probability on spot checks). Most of the gap to the published score is already present at 8-bit — sample noise, harness differences, and Arts & Human Taste in particular — not quantization. Latency is per request on Apple Silicon (median request ~1k tokens); MLX prefill is ~2k tok/s for the 9B, slower for the 27B.
Method: Decision Index 0.2.1 kit (suite rebuilt byte-identical), stratified
2,000-request sample across all 44 benchmarks (suite sample --n 2000), engine = clef_mlx.py with
max_length=16384 and no truncation (over-length requests are refused and count as wrong; 15 of 2,000,
mostly BRIGHT). The index is computed from each benchmark's native metric on the sampled rows with the kit's
chance correction and weights; HLE and iSarcasmEval are set to 0 to match how Cloudflare's published run is
scored. With ~40 rows per benchmark, per-benchmark numbers are noisy (±10+ pts) — only the index is meaningful.
License
Apache-2.0, following Cloudflare/clef.
- Downloads last month
- 110
4-bit