Minima

A strict W1.58A8 adaptation of LiquidAI/LFM2.5-Encoder-350M, built with SSHDotCodes/minima.

  • Logical matrix values: {-1, 0, +1} (1.585 bits)
  • Physical artifact format: I2_S, four trits per byte
  • Dynamic int8 activations
  • Group size 128, no recovery adapters
  • Full 8,192-token encoder context
  • Packed weight file: 89.9 MB (94,298,568 bytes)
  • Release status: release candidate (six-task gated retention 85.45% vs FP32)

Use

pip install "minima-lfm @ git+https://github.com/SSHDotCodes/minima.git"
from minima import MinimaModel

model = MinimaModel.from_pretrained("ProCreations/minima", device="cpu")
outputs = model(input_ids=input_ids, attention_mask=attention_mask)

CPU inference defaults to a one-time FBGEMM dynamic-int8 packing of each ternary matrix. Set MINIMA_CPU_BACKEND=i2s for the direct packed 2-bit AVX2/ARM NEON kernel.

Encoder quality

The six-task downstream gate compares matched 800-step fine-tunes. Packed ternary matrices stay frozen; only the task head and non-matrix parameters adapt. Required relative mean: >= 96%.

Task FP32 Minima Capped retention
SST2 0.79817 0.75459 94.54%
QNLI 0.60608 0.59491 98.16%
MNLI 0.40621 0.38074 93.73%
MRPC 0.81694 0.81873 100.00%
STSB 0.54812 0.40660 74.18%
COLA 0.11382 0.05927 52.07%
Mean 85.45%

License

The weights remain subject to the LFM Open License v1.0 shipped in this repository. The Minima runtime code is MIT licensed.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
91.4M params
Tensor type
F16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ProCreations/minima

Finetuned
(27)
this model