Instructions to use wesleysimplicio/Simplicio-27B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use wesleysimplicio/Simplicio-27B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf wesleysimplicio/Simplicio-27B:BF16 # Run inference directly in the terminal: llama cli -hf wesleysimplicio/Simplicio-27B:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf wesleysimplicio/Simplicio-27B:BF16 # Run inference directly in the terminal: llama cli -hf wesleysimplicio/Simplicio-27B:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf wesleysimplicio/Simplicio-27B:BF16 # Run inference directly in the terminal: ./llama-cli -hf wesleysimplicio/Simplicio-27B:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf wesleysimplicio/Simplicio-27B:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf wesleysimplicio/Simplicio-27B:BF16
Use Docker
docker model run hf.co/wesleysimplicio/Simplicio-27B:BF16
- LM Studio
- Jan
- vLLM
How to use wesleysimplicio/Simplicio-27B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "wesleysimplicio/Simplicio-27B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "wesleysimplicio/Simplicio-27B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/wesleysimplicio/Simplicio-27B:BF16
- Ollama
How to use wesleysimplicio/Simplicio-27B with Ollama:
ollama run hf.co/wesleysimplicio/Simplicio-27B:BF16
- Unsloth Desktop
- Pi
How to use wesleysimplicio/Simplicio-27B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wesleysimplicio/Simplicio-27B:BF16
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "wesleysimplicio/Simplicio-27B:BF16" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use wesleysimplicio/Simplicio-27B with Docker Model Runner:
docker model run hf.co/wesleysimplicio/Simplicio-27B:BF16
- Lemonade
How to use wesleysimplicio/Simplicio-27B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull wesleysimplicio/Simplicio-27B:BF16
Run and chat with the model
lemonade run user.Simplicio-27B-BF16
List all available models
lemonade list
- Hermes Agent
How to use wesleysimplicio/Simplicio-27B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wesleysimplicio/Simplicio-27B:BF16
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default wesleysimplicio/Simplicio-27B:BF16
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use wesleysimplicio/Simplicio-27B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf wesleysimplicio/Simplicio-27B:BF16
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "wesleysimplicio/Simplicio-27B:BF16" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Simplicio 27B
A Qwen3.8-27B fine-tune that answers code-change requests with SEARCH/REPLACE patches instead of whole files.
Overview
Simplicio 27B is a LoRA fine-tune of Qwen3.8-27B by Wesley Simplicio at SimpleTI. It answers a code-change request in five tagged phases (<orient>, <plan>, <patch>, <validate>, <deliver>). The <patch> phase holds SEARCH/REPLACE blocks that touch only the lines that change.
- Base model: Qwen3.8-27B, 27.36B parameters, 64 layers that alternate linear attention (DeltaNet) and full attention in a 3:1 pattern.
- Files on Hugging Face: the LoRA adapter (0.64 GB), the merged BF16 checkpoint (18 shards, 55.6 GB), and a GGUF Q4_K_M (16.8 GB) with the base model's vision projector (0.93 GB).
- Status: research release. It passes 46.7% of the project's own 120 held-out tasks. It has not been scored on public benchmarks.
Results
Held-out set (120 tasks)
The full BF16 model ran on Google Colab G4 (NVIDIA RTX PRO 6000 Blackwell, 95 GB) against 120 tasks that were not used in training (data/unseen_eval_120.json). Each task got one attempt at temperature 0, and generation stopped at </deliver>. A task passes when its unit test passes after the patch is applied. Aggregate results: benchmarks/live_colab_g4_bf16_n120.json.
| Metric | Result |
|---|---|
| Unit test passes | 56/120 (46.7%) |
| SEARCH block found, patch applied | 104/120 (86.7%) |
| Patched file parses (AST) | 104/120 (86.7%) |
| No call to a nonexistent API | 120/120 |
| Output tokens per task, mean / max | 69.9 / 103 |
| Wall time for all 120 tasks | 438 s |
| Category | Tasks | Unit test passes | Patch applied |
|---|---|---|---|
| Surgical diff and AST precision | 30 | 17 | 20 |
| Edge-case correctness | 30 | 6 | 27 |
| Nonexistent and deprecated API traps | 30 | 27 | 30 |
| Adversarial and out-of-distribution | 30 | 6 | 27 |
The base model has not been run under this protocol yet, so the gain from fine-tuning is not measured.
Withdrawn numbers
Earlier versions of this card reported 96.5% accuracy, 116 of 120 tasks passed, 480 tokens per task, a "Top 12" leaderboard and per-token prices. None of these came from running the model on those tasks:
- 116/120 and its McNemar test come from
benchmarks/prove_benchmark_120.py, which simulates model outputs. - 480 tokens per task comes from a 3-task smoke test on an A100 (
benchmarks/empirical_a100_results.json). - 96.5% and the leaderboard rows are hard-coded in
benchmarks/compare_top10_2026.py. - The model is not listed on OpenRouter and has no public per-token price.
The held-out run above replaces them.
Public coding-agent benchmarks
Simplicio 27B has not been evaluated on SWE-bench, the Aider benchmark, Terminal-Bench or the Artificial Analysis Coding Agent Index. For scale, these are published Coding Agent Index v1.5 results (Artificial Analysis, retrieved 5 October 2026):
| Agent | Coding Agent Index | DeepSWE v1.1 | Terminal-Bench 4.0 | SWE-Atlas-QnA | Cost per task |
|---|---|---|---|---|---|
| Claude Opus 5.5 (max) | 66 | 68% | 63% | 66% | $13.04 |
| Claude Sonnet 5.5 (max) | 68 | 72% | 66% | 67% | $14.19 |
| GPT-6.1 Sol (xhigh) | 63 | 73% | 55% | 61% | $1.04 |
| Simplicio 27B | not evaluated | – | – | – | – |
Quick start
Ollama
ollama run wesleysimplicio/simplicio-27b
The latest tag holds the Q4_K_M GGUF and the vision projector. It uses temperature 0.2 and a 32,768-token context, and it stops at <|im_end|> and </deliver>.
The installer installs Ollama if it is missing, then runs the model:
curl -fsSL https://raw.githubusercontent.com/simpletibr/simplicio-27b/main/install.sh | bash
vLLM (OpenAI-compatible server with tool calls)
git clone https://github.com/simpletibr/simplicio-27b
cd simplicio-27b
./deploy/serve_vllm.sh wesleysimplicio/Simplicio-27B 8000
This serves the merged BF16 checkpoint. The weights alone take 55.6 GB; the evaluation above ran on a 95 GB GPU. The script sets:
--max-model-len 40960, defined once indeploy/context.env: a measured 31,692-token OpenCode prompt plus 4,096 output tokens.--chat-templatewithdeploy/chat_template_chatml.jinja. It prefills<think>so that--reasoning-parser qwen3moves reasoning out ofcontent.--enable-auto-tool-choice --tool-call-parser simplicio, usingdeploy/simplicio_tool_parser.py. It turns<tool><name>…</name><params>…</params></tool>into a singletool_callsentry.--served-model-name simplicio-27b simpleti/simplicio-27b.
On a smaller GPU, Simplicio_27B_Serve_Colab.ipynb serves the 4-bit base with the LoRA adapter on Colab.
OpenCode
Add the vLLM server to opencode.json as an OpenAI-compatible provider:
{
"$schema": "https://opencode.ai/config.json",
"provider": {
"simplicio": {
"npm": "@ai-sdk/openai-compatible",
"name": "Simplicio 27B",
"options": { "baseURL": "http://localhost:8000/v1" },
"models": { "simplicio-27b": { "name": "Simplicio 27B" } }
}
}
}
opencode -m simplicio/simplicio-27b
OpenCode works through tool calls, so point it at the vLLM server. The Ollama template does not declare tools.
Python (Unsloth)
This loads the adapter on its 4-bit base, the same way Simplicio_27B_Merge_Colab.ipynb does:
from unsloth import FastLanguageModel
model, tokenizer = FastLanguageModel.from_pretrained(
model_name="wesleysimplicio/Simplicio-27B", # adapter; the base comes from adapter_config.json
max_seq_length=16384,
load_in_4bit=True,
)
FastLanguageModel.for_inference(model)
messages = [{"role": "user", "content": "In api/schemas/user.py, accept tax_id with punctuation such as 123.456.789-00."}]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to("cuda")
output = model.generate(inputs, max_new_tokens=1024, do_sample=False)
print(tokenizer.decode(output[0][inputs.shape[1]:], skip_special_tokens=True))
Output format
Each phase is a list of numbered points. This example is shortened and translated from the first training example:
<simplicio_loop>
<orient>
[Point 1: Root] Root confirmed at /workspace/api-gateway (pyproject.toml found).
[Point 3: Type signatures] UserCreate.tax_id: str = Field(..., min_length=11, max_length=14).
…
</orient>
<plan>
[Point 11: Atomic steps] Step 1: add a field_validator to the schema. Step 2: run tests/test_users.py.
…
</plan>
<patch>
<<<< SEARCH
tax_id: str = Field(..., min_length=11, max_length=14)
====
tax_id: str = Field(..., min_length=11, max_length=11)
@field_validator("tax_id", mode="before")
@classmethod
def sanitize_tax_id(cls, v: str) -> str:
cleaned = re.sub(r"\D", "", v)
if len(cleaned) != 11:
raise ValueError("tax_id must have exactly 11 digits")
return cleaned
>>>> REPLACE
</patch>
<validate>
[Point 32: Targeted tests] pytest tests/test_users.py -k "tax_id" -> 2 passed
…
</validate>
<deliver>
…
</deliver>
</simplicio_loop>
The SEARCH/REPLACE markers are four characters long (<<<<, ====, >>>>), not the seven that Git and Aider use. To apply a patch, find the SEARCH text verbatim in the file and replace it.
Training
The published adapter was produced by Simplicio_27B_Training_Colab.ipynb. The settings below come from that notebook and the published adapter_config.json.
| Setting | Value |
|---|---|
| Method | QLoRA with Unsloth; base loaded in 4-bit |
| LoRA | r 32, alpha 32, dropout 0, on the q, k, v, o, gate, up and down projections of every layer |
| Steps | 120 steps, batch size 1, gradient accumulation 8 |
| Optimizer | AdamW 8-bit, learning rate 2e-4, cosine schedule, 10 warmup steps, weight decay 0.01 |
| Sequence length | 4,096 |
| Loss | whole sequence, prompt included |
| Data | 101 examples in Portuguese: 1 written by hand and 100 generated from a short list of stack and task templates |
| Hardware | Google Colab A100 (40 GB) |
train_simplicio_27b.py is a script version with extra options: freezing the bottom layers, attention-only LoRA, and registering the phase tags as special tokens. The published adapter used none of them, and its tokenizer has no added tokens. generate_dataset.py writes data/simplicio_loop_50pts_train.jsonl (80 examples) and data/simplicio_loop_50pts_val.jsonl (15 examples).
Limitations
- It passes 46.7% of the held-out tasks, and only 6 of 30 in both the edge-case and the adversarial categories.
- The training set is small (101 examples), templated and in Portuguese. The model follows the format more reliably than it solves the task.
- In the training examples,
<validate>and<deliver>contain written-out results such as "2 passed" or "COMMIT_READY". The model writes these without running anything. Treat them as claims and run your own tests. - It has not been compared with the base model under the same protocol, and it has not been run on public benchmarks.
- Aider: the patch markers differ from Aider's edit format, and Aider has not been tested.
- Vision: the GGUF ships the base model's vision projector. Training was text-only, and image input has not been evaluated.
Repository
| Path | Contents |
|---|---|
Simplicio_27B_Training_Colab.ipynb |
Training run that produced the adapter |
Simplicio_27B_Merge_Colab.ipynb |
Merges the adapter into 16-bit weights and exports the GGUF Q4_K_M |
Simplicio_27B_Serve_Colab.ipynb, deploy/ |
vLLM serving, chat template, tool parser, context length, Ollama Modelfile |
data/unseen_eval_120.json |
The 120 held-out tasks |
benchmarks/live_colab_g4_bf16_n120.json |
The results above |
tests/ |
Tests for the serving code: python -m pytest tests |
Citation
@misc{simplicio27b2026,
author = {Simplicio, Wesley},
title = {Simplicio 27B: a Qwen3.8-27B fine-tune for SEARCH/REPLACE code patches},
year = {2026},
publisher = {SimpleTI},
howpublished = {\url{https://huggingface.co/wesleysimplicio/Simplicio-27B}}
}
Links
- Product page: simpleti.com.br/simplicio-27b/
- Weights: huggingface.co/wesleysimplicio/Simplicio-27B
- Ollama: ollama.com/wesleysimplicio/simplicio-27b
- Source: github.com/simpletibr/simplicio-27b
- Base model: Qwen/Qwen3.8-27B (Apache 2.0)
- Fine-tuning: Unsloth
- Downloads last month
- 918