policyfx-nli

Fine-tuned version of tasksource/ModernBERT-base-nli for natural language inference on Congressional Research Service (CRS) policy reports.

Intended Use

Classifies whether a policy passage entails, contradicts, or is neutral with respect to a stakeholder impact statement.

Input Format

premise    = "The program provides monthly cash payments of $500 to eligible low-income families."
hypothesis = "low-income families: Low-income families receive monthly cash payments of $500 through the program."

The hypothesis follows the format {stakeholder}: {effect statement}. The premise–hypothesis pair is truncated to 2048 tokens combined.

Usage

from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
from scipy.special import softmax

model_id = "JoshuaAshkinaze/policyfx-nli"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

LABEL_NAMES = ["contradiction", "entailment", "neutral"]

premise    = "The bill will raise taxes."
hypothesis = "Taxpayers: Individual taxpayers will have less disposable income."

inputs = tokenizer(premise, hypothesis, return_tensors="pt",
                   truncation=True, max_length=2048)
with torch.no_grad():
    logits = model(**inputs).logits.numpy()

probs = softmax(logits, axis=-1)[0]
pred  = {label: round(float(p), 4) for label, p in zip(LABEL_NAMES, probs)}
print(pred)

Note on label order: config.id2label is set to {0: "contradiction", 1: "entailment", 2: "neutral"}. Do not remap from the base model's label config.

Training

Data: 61,599 (premise, stakeholder-effect hypothesis, label) triples generated from ~3,000 CRS reports, covering topics from housing and healthcare to defense and immigration, balanced across labels (20,533 per class). Train/val/test split is report-level to prevent leakage.

Hard example filtering: The base NLI model scored all examples; only those with P(true_label) < 0.80 were retained. The final hard dataset contains 19,965 examples (6,655 per class). This removes trivially easy examples and focuses training on genuine domain difficulty.

Spurious correlation check: A TF-IDF classifier trained on hypotheses alone achieves macro-F1 = 0.34 (≈ chance), confirming that labels are not recoverable from surface form.

Splits: 14,085 train, 3,006 validation, 2,874 test.

Hyperparameter search: We trained four configurations (learning rate 1e-5 or 2e-5, crossed with 1 or 2 epochs), evaluating validation macro-F1 after each epoch and keeping the best epoch per run. We selected the configuration with the highest validation macro-F1; the test set was not used for selection. We then retrained that configuration from the base model on train+val.

Learning rate Max epochs Best epoch Val macro-F1
1e-5 1 1 0.9567
1e-5 2 2 0.9610
2e-5 1 1 0.9640
2e-5 2 2 0.9647

Hyperparameters (best config, retrained on train+val):

Parameter Value
Base model tasksource/ModernBERT-base-nli
Learning rate 2e-5
Epochs 2
Batch size 16
Weight decay 0.01
Max length 2048
Optimizer AdamW (fused)
Precision bfloat16 + tf32
Seed 42
Parameters 149,607,171

Evaluation

Evaluated on a held-out test set of 2,874 balanced examples (958 per class).

Metric Score
Macro-F1 0.9718
ROC-AUC (OvR) 0.9968
Brier Score 0.0493
ECE 0.0211

Per-class F1:

Label Precision Recall F1
Contradiction 0.956 0.961 0.959
Entailment 0.962 0.955 0.959
Neutral 0.997 0.999 0.998

Limitations

  • Domain specificity: Fine-tuned on CRS policy text with a specific hypothesis format ({stakeholder}: {effect}). Performance on general NLI benchmarks (SNLI, MNLI, ANLI) is lower than the zero-shot base model — this is an expected cost of domain adaptation.
  • Hypothesis format: The model expects the stakeholder-prefixed hypothesis format. Inputs that deviate significantly from this format may produce unreliable probabilities.
  • English only.

Citation

@misc{ashkinaze2026policyfx,
  author    = {Ashkinaze, Joshua},
  title     = {policyfx-nli: ModernBERT-NLI fine-tuned on Congressional Research Service policy reports},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/JoshuaAshkinaze/policyfx-nli}
}
Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JoshuaAshkinaze/policyfx-nli

Finetuned
(8)
this model