Escarda-86M-Identity

Identity-tuned chat variant of Escarda-86M (SFT epoch 3). ~86M SpikeWhaleLM (JEPA + HRM refine), ChatML-aware SpikeTokenizer. It knows it is "Escarda" and answers like a clean assistant.

Usage

Custom architecture + tokenizer. trust_remote_code=True:

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("Quazim0t0/Escarda-86M-Identity", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("Quazim0t0/Escarda-86M-Identity", trust_remote_code=True)

ChatML (<|im_start|>role\nโ€ฆ<|im_end|>). Generation starts after a trailing <|im_start|>assistant\n.

Architecture

SpikeWhaleLM (~86M, 16 layers, hidden 640, 4096 context, 16,512 vocab, tied embeddings): MLA (LoRA-rank-128 Q/O, decoupled RoPE-16 + NoPE-48, multi-query, QK-norm), engram n-gram memory, ร—2 hash-lookup, hyper-connections, HRM refine, MTP training head, and JEPA. Escarda uses both (use_hrm_refine=True, use_jepa=True).

Tokenizer

SpikeTokenizer - byte-level length-max (greedy longest-match), 16,512 vocab, ChatML atomic specials. PreTrainedTokenizer. Load with AutoTokenizer + trust_remote_code.

Evaluation

Zero-shot, full val/test splits (acc = raw continuation log-likelihood, acc_norm = byte-length-normalized).

Task acc acc_norm
ARC-Easy 0.3262 0.3380
ARC-Challenge 0.2048 0.2415
HellaSwag 0.2785 0.2818
WinoGrande 0.5020 -
PIQA 0.5539 0.5462
OpenBookQA 0.1360 0.2440
BoolQ 0.4174 -

ArithMark-2.0 (AxiomicLabs)

  • official metric is raw acc: 0.3628 (strongest of the Escarda family).

Language modeling: WikiText-2 byte-ppl โ†“ 2.7062 ยท BLiMP โ†‘ 0.7133.

Live demo: Escarda-86M-Chat Space.

Citation

If you use this model, please cite:

@misc{escarda86midentity,
  title        = {Escarda-86M-Identity: A ~86M-parameter SpikeWhaleLM},
  author       = {Dean Byrne (Quazim0t0)},
  year         = {2026},
  howpublished = {HuggingFace, \url{https://huggingface.co/Quazim0t0/Escarda-86M-Identity}},
  note         = {Quazim0t0/Escarda-86M-Identity}
}

Update: format-blended SFT on the engram-repaired base

This revision applies the (behavior-preserving) engram repair, then a short instruction/format SFT on a 60/25/15 blend of HuggingFaceTB/smoltalk, GSM8K-train (with '#### N' reasoning), and MMLU-style ('Answer: ') examples -- so chat fluency improves while the benchmark output-formats are preserved rather than overwritten. Held-out (test-split) before->after:

MMLU acc 0.188->0.284, format 0.700->0.958; GSM8K '####' 0.530->0.735

Note: these are fluency + output-format gains. Benchmark accuracy remains near the floor for a model this size -- the SFT does not add reasoning ability.

Downloads last month
755
Safetensors
Model size
97.3M params
Tensor type
F32
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Quazim0t0/Escarda-86M-Identity

Finetunes
2 models

Spaces using Quazim0t0/Escarda-86M-Identity 3