Instructions to use Quazim0t0/Escarda-86M-Identity with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Quazim0t0/Escarda-86M-Identity with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Quazim0t0/Escarda-86M-Identity", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("Quazim0t0/Escarda-86M-Identity", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Quazim0t0/Escarda-86M-Identity with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Quazim0t0/Escarda-86M-Identity" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Escarda-86M-Identity", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Quazim0t0/Escarda-86M-Identity
- SGLang
How to use Quazim0t0/Escarda-86M-Identity with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Quazim0t0/Escarda-86M-Identity" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Escarda-86M-Identity", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Quazim0t0/Escarda-86M-Identity" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Quazim0t0/Escarda-86M-Identity", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Quazim0t0/Escarda-86M-Identity with Docker Model Runner:
docker model run hf.co/Quazim0t0/Escarda-86M-Identity
Escarda-86M-Identity
Identity-tuned chat variant of Escarda-86M
(SFT epoch 3). ~86M SpikeWhaleLM (JEPA + HRM refine), ChatML-aware SpikeTokenizer.
It knows it is "Escarda" and answers like a clean assistant.
Usage
Custom architecture + tokenizer. trust_remote_code=True:
from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Quazim0t0/Escarda-86M-Identity", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Quazim0t0/Escarda-86M-Identity", trust_remote_code=True)
ChatML (<|im_start|>role\nโฆ<|im_end|>). Generation starts after a trailing
<|im_start|>assistant\n.
Architecture
SpikeWhaleLM (~86M, 16 layers, hidden 640, 4096 context, 16,512 vocab, tied
embeddings): MLA (LoRA-rank-128 Q/O, decoupled RoPE-16 + NoPE-48, multi-query,
QK-norm), engram n-gram memory, ร2 hash-lookup, hyper-connections, HRM refine,
MTP training head, and JEPA. Escarda uses both (use_hrm_refine=True,
use_jepa=True).
Tokenizer
SpikeTokenizer - byte-level length-max (greedy longest-match), 16,512 vocab,
ChatML atomic specials. PreTrainedTokenizer. Load with AutoTokenizer +
trust_remote_code.
Evaluation
Zero-shot, full val/test splits (acc = raw continuation log-likelihood,
acc_norm = byte-length-normalized).
| Task | acc | acc_norm |
|---|---|---|
| ARC-Easy | 0.3262 | 0.3380 |
| ARC-Challenge | 0.2048 | 0.2415 |
| HellaSwag | 0.2785 | 0.2818 |
| WinoGrande | 0.5020 | - |
| PIQA | 0.5539 | 0.5462 |
| OpenBookQA | 0.1360 | 0.2440 |
| BoolQ | 0.4174 | - |
ArithMark-2.0 (AxiomicLabs)
- official metric is raw
acc: 0.3628 (strongest of the Escarda family).
Language modeling: WikiText-2 byte-ppl โ 2.7062 ยท BLiMP โ 0.7133.
Live demo: Escarda-86M-Chat Space.
Citation
If you use this model, please cite:
@misc{escarda86midentity,
title = {Escarda-86M-Identity: A ~86M-parameter SpikeWhaleLM},
author = {Dean Byrne (Quazim0t0)},
year = {2026},
howpublished = {HuggingFace, \url{https://huggingface.co/Quazim0t0/Escarda-86M-Identity}},
note = {Quazim0t0/Escarda-86M-Identity}
}
Update: format-blended SFT on the engram-repaired base
This revision applies the (behavior-preserving) engram repair, then a short instruction/format SFT on a 60/25/15 blend of HuggingFaceTB/smoltalk, GSM8K-train (with '#### N' reasoning), and MMLU-style ('Answer: ') examples -- so chat fluency improves while the benchmark output-formats are preserved rather than overwritten. Held-out (test-split) before->after:
MMLU acc 0.188->0.284, format 0.700->0.958; GSM8K '####' 0.530->0.735
Note: these are fluency + output-format gains. Benchmark accuracy remains near the floor for a model this size -- the SFT does not add reasoning ability.
- Downloads last month
- 755