Instructions to use RandomUs666/ecg-interface-study with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use RandomUs666/ecg-interface-study with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("RandomUs666/ecg-interface-study") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use RandomUs666/ecg-interface-study with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "RandomUs666/ecg-interface-study" --prompt "Once upon a time"
- Atomic Chat
ECG interface study β four LoRA adapters, one controlled comparison
Research prototype. Not a medical device, not for clinical use, not validated for any patient population. These adapters exist to answer a methodological question, not to read an ECG for you.
Access is gated with automatic approval: the repository is open to read, and downloading asks you for an account and an acknowledgement of the line above. Approval is immediate; the gate is there so that nobody ends up with clinical-looking weights without having been told what they are.
This repository holds the four LoRA adapters from a study asking one question: how much does the way an ECG is handed to a language model matter, compared with which language model you pick?
The answer, measured on one identical set of 1,000 patient recordings: the interface moved macro-averaged F1 by about 0.55; swapping model family at a fixed interface moved it by 0.085.
The comparison
All four adapters were trained on the same 1,000 Chapman-Shaoxing/Ningbo recordings, the same 11 diagnostic classes, the same LoRA shape (rank 16, scale 2.0, dropout 0, batch 1, seed 42).
One caveat about the third row, because the naming invites a wrong reading: measures-qwen35-2b-8layer
is not a smaller model. It is the same Qwen3.5-2B, whole, with LoRA attached to only the last 8 of
its 24 transformer blocks instead of all of them (num_layers: 8 in mlx_lm.lora means "convert this
many blocks, counting from the last", not "truncate the network"). The choice was originally a speed
lever. What its 0.2909 measures is the cost of letting only a third of the network adapt β a property
of the training setup, not of model size. They are
evaluated on population V: 1,000 held-out validation recordings, one shared scoring protocol,
95 % confidence intervals from a 10,000-resample record-level bootstrap.
| folder | input the model reads | macro-F1 [95 % CI] | AUROC |
|---|---|---|---|
measures-gemma-4-e2b/ |
32 clinical measurements as one line of text | 0.5960 [0.5565, 0.6281] | 0.9020 |
measures-qwen35-2b-full/ |
the same measurements | 0.5112 [0.4729, 0.5459] | 0.9262 |
measures-qwen35-2b-8layer/ |
the same measurements, but only the last 8 of 24 blocks carry adapters | 0.2909 [0.2640, 0.3167] | 0.6563 |
tokens-gemma-4-e2b/ |
400 discrete codes from a learned waveform codebook | 0.0436 [0.0393, 0.0475] | β |
Reference points measured on the same 1,000 records, not in this repository: gradient boosting on the same 32 measurements 0.6407 [0.5983, 0.6746]; a 476,819-parameter 1-D convolutional network on the continuous waveform 0.5881 [0.5489, 0.6214]; predicting the training prevalence and nothing else 0.1853.
The first and last rows are the same model family and the same LoRA shape on the same recordings. Only the interface differs β and with it two training settings that were not held constant, stated plainly because they bound what the comparison proves:
tokens-gemma-4-e2b |
measures-gemma-4-e2b |
|
|---|---|---|
| representation | 400 discrete codes | 32 clinical measurements as text |
| supervised decisions per example | ~1.7 | 11 |
| training length | 1 epoch (iters: 1000) |
10 epochs (iters: 10000), read out at epoch 8 |
| learning rate | 0.0001 | 0.00005 |
max_seq_length |
2048 | 512 |
Because four things differ, the 0.55 swing is attributed to the interface as a whole β the representation and the form of the question together β and to neither half alone. A ten-epoch re-run of the token configuration was monitored per epoch and the collapse was still in place when it was stopped at epoch 5.7: directional evidence that training length alone does not explain the gap, not a completed control.
What the token adapter is for
tokens-gemma-4-e2b/ scores 0.0436, far below the 0.1853 you get by predicting class prevalence.
It is published deliberately, as the negative half of the comparison. It emits one constant answer for
every input β and not even the answer the class frequencies would recommend: the most common class in
the corpus is never emitted at all.
Three controls locate that failure in the interface rather than the architecture:
- the identical recipe, given 16 examples and 100 epochs, reproduces all 16 label sets exactly;
- a linear probe on the very same discrete codes reaches macro-F1 0.3947 (n = 4,359), so the class information survived the encoding;
- classical models on those codes plateau at 0.4703 at this training size, about 0.10 below what the continuous signal supports β so the representation was not innocent either.
Do not read this as "language models cannot interpret ECGs." That was never tested at an adequate supervision budget. What was measured is that ~1.7 scored decisions per example is not enough, against the ~330,000 the paired convolutional network received on the same 1,000 recordings.
Usage
These are LoRA adapters for MLX. They need their base model.
pip install mlx-lm
huggingface-cli download RandomUs666/ecg-interface-study --local-dir ecg-interface-study
python -m mlx_lm.generate \
--model mlx-community/gemma-4-e2b-it-bf16 \
--adapter-path ecg-interface-study/measures-gemma-4-e2b \
--prompt "$(cat ecg-interface-study/prompts/formA_slots_system.txt)
FC 148 bpm | RR 405Β±85 ms | RMSSD 116 ms | battiti 25, prematuri 5 (larghi 5, stretti 0) | P presente (0.54 mV) | PR 290 ms | QRS 140 ms | QT 300 ms | QTc 487 ms | asse 50 | R V1 0.09 mV | S V1 1.30 mV" \
--max-tokens 32 --temp 0.0
Expected output: exactly 11 SI/NO tokens, space-separated, one per class in the order given in
the system prompt. Greedy decoding (--temp 0.0); the reported numbers come from greedy decoding and
a frozen parser that counts unparseable output as an empty prediction set.
The prompt is not decoration β it is the interface, and the interface is the finding. The system
prompts are in prompts/, verbatim from the training data, in Italian, and the adapters expect them
as-is. formA_slots_system.txt drives the three measures-* adapters; vq_tokens_system.txt drives
tokens-gemma-4-e2b, whose output format is instead a JSON object {"labels": [...]}.
Each folder ships every intermediate checkpoint (0001000_adapters.safetensors β¦), so the training
trajectory can be inspected, plus adapters.safetensors, the final one. The reported number for
measures-gemma-4-e2b is the epoch-8 checkpoint, 0008000_adapters.safetensors, not the final one.
Epoch 8 was selected in advance on a separate 512-record monitor; it is not established as better than
epoch 7 (difference +0.0130, CI [β0.0217, +0.0476]).
Classes
Eleven multi-label classes, each with at least 200 training positives, covering 95.0 % of the corpus:
ATRIAL_FIB_FLUTTER, AVB_FIRST_DEGREE, LEFT_CONDUCTION_ABNORMALITY, NORMAL_ECG, PAC, PVC,
RBBB, SINUS_BRADYCARDIA, SINUS_RHYTHM, SINUS_TACHYCARDIA, SUPRAVENTRICULAR_TACHYCARDIA.
Data and protocol
- Corpus: Chapman-Shaoxing/Ningbo, 45,152 physician-reviewed recordings at 500 Hz, via PhysioNet.
- Splits: 80/10/10 by patient, seed 42, iteratively stratified, each split frozen with a SHA-256 hash. The split key is the signal-hash duplicate group, not the record id: hashing every waveform found 87 groups of byte-identical duplicates in a corpus documented as one recording per patient, 32 of which straddled a naive split.
- Training budget: 1,000 recordings, a hardware ceiling rather than a choice β everything ran on one laptop (Apple M4 Pro, 48 GB unified memory) under a 37 GiB working-set cap.
- Thresholds: the language models emit hard labels, so no per-class thresholds are fitted for them.
- Test split: opened exactly once, for other systems. The measurement arm has no test-set number at all, and none should be invented for it. Every figure above is validation.
results/v2_results.json is the primary artefact behind the table: per-system scores, per-class F1,
the nine paired comparisons with bootstrap intervals and exact McNemar tests.
Honest limits
- One corpus, one run, one seed, one machine. No repeated seeds, no multiple-comparison correction across the campaign's 28 pre-registered pairs, nine of which are not established.
- 1,000 of 34,658 available recordings. That cut alone costs the paired convolutional network 0.2485 macro-F1, so every number here sits far below what the data supports.
- Representation and supervision were never crossed, nor were training length and learning rate held constant (table above). The decomposition of the 0.55 swing is not available from this study.
- Single-beat events are where every aggregate interface fails: the best adapter here scores 0.192 on premature ventricular complexes against 0.338 for gradient boosting on the same features.
- Prompts and outputs are in Italian. The adapters were never trained or tested in any other language.
- The models are not calibrated: they emit labels, not probabilities. AUROC in the table comes from a separate scoring path and is reported beside F1 because the two disagree β checkpoints spanning AUROC 0.90β0.93 span 0.51β0.60 in thresholded F1.
What this study does not claim
- Not that language models cannot read ECGs.
- Not that a language model beats a purpose-built convolutional network. On the same 1,000 records the difference is +0.0079 with a confidence interval of [β0.036, +0.053] and a McNemar p of 0.67: they are indistinguishable. An earlier version of this work claimed the win; it came from a smaller evaluation (n = 256) in which only the language models were granted fitted thresholds, and it is retracted.
- Not that the language model beats gradient boosting β boosting is ahead by 0.0447, interval clear of zero.
- Nothing clinical. No sensitivity, specificity or safety claim is made or implied.
Licensing
Mixed, and inherited β check before redistributing:
measures-gemma-4-e2b/andtokens-gemma-4-e2b/adapt Gemma and are subject to the Gemma Terms of Use and its Prohibited Use Policy.measures-qwen35-2b-full/andmeasures-qwen35-2b-8layer/adapt Qwen; see that model's licence.- The training corpus is Chapman-Shaoxing/Ningbo, CC BY 4.0, obtained from PhysioNet. No patient data, waveform, or derived recording is included in this repository β only adapter weights, the prompt templates and aggregate results.
- Everything authored here (documentation, prompts, results JSON) : CC BY 4.0.
Citation
An academic write-up of this study exists but is not yet public. Until then, cite this repository by its URL and commit.
Quantized