policyfx-nli
Fine-tuned version of tasksource/ModernBERT-base-nli for natural language inference on Congressional Research Service (CRS) policy reports.
Intended Use
Classifies whether a policy passage entails, contradicts, or is neutral with respect to a stakeholder impact statement.
Input Format
premise = "The program provides monthly cash payments of $500 to eligible low-income families."
hypothesis = "low-income families: Low-income families receive monthly cash payments of $500 through the program."
The hypothesis follows the format {stakeholder}: {effect statement}. The premise–hypothesis pair is truncated to 2048 tokens combined.
Usage
from transformers import AutoTokenizer, AutoModelForSequenceClassification
import torch
from scipy.special import softmax
model_id = "JoshuaAshkinaze/policyfx-nli"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()
LABEL_NAMES = ["contradiction", "entailment", "neutral"]
premise = "The bill will raise taxes."
hypothesis = "Taxpayers: Individual taxpayers will have less disposable income."
inputs = tokenizer(premise, hypothesis, return_tensors="pt",
truncation=True, max_length=2048)
with torch.no_grad():
logits = model(**inputs).logits.numpy()
probs = softmax(logits, axis=-1)[0]
pred = {label: round(float(p), 4) for label, p in zip(LABEL_NAMES, probs)}
print(pred)
Note on label order: config.id2label is set to {0: "contradiction", 1: "entailment", 2: "neutral"}. Do not remap from the base model's label config.
Training
Data: 61,599 (premise, stakeholder-effect hypothesis, label) triples generated from ~3,000 CRS reports, covering topics from housing and healthcare to defense and immigration, balanced across labels (20,533 per class). Train/val/test split is report-level to prevent leakage.
Hard example filtering: The base NLI model scored all examples; only those with P(true_label) < 0.80 were retained. The final hard dataset contains 19,965 examples (6,655 per class). This removes trivially easy examples and focuses training on genuine domain difficulty.
Spurious correlation check: A TF-IDF classifier trained on hypotheses alone achieves macro-F1 = 0.34 (≈ chance), confirming that labels are not recoverable from surface form.
Splits: 14,085 train, 3,006 validation, 2,874 test.
Hyperparameter search: We trained four configurations (learning rate 1e-5 or 2e-5, crossed with 1 or 2 epochs), evaluating validation macro-F1 after each epoch and keeping the best epoch per run. We selected the configuration with the highest validation macro-F1; the test set was not used for selection. We then retrained that configuration from the base model on train+val.
| Learning rate | Max epochs | Best epoch | Val macro-F1 |
|---|---|---|---|
| 1e-5 | 1 | 1 | 0.9567 |
| 1e-5 | 2 | 2 | 0.9610 |
| 2e-5 | 1 | 1 | 0.9640 |
| 2e-5 | 2 | 2 | 0.9647 |
Hyperparameters (best config, retrained on train+val):
| Parameter | Value |
|---|---|
| Base model | tasksource/ModernBERT-base-nli |
| Learning rate | 2e-5 |
| Epochs | 2 |
| Batch size | 16 |
| Weight decay | 0.01 |
| Max length | 2048 |
| Optimizer | AdamW (fused) |
| Precision | bfloat16 + tf32 |
| Seed | 42 |
| Parameters | 149,607,171 |
Evaluation
Evaluated on a held-out test set of 2,874 balanced examples (958 per class).
| Metric | Score |
|---|---|
| Macro-F1 | 0.9718 |
| ROC-AUC (OvR) | 0.9968 |
| Brier Score | 0.0493 |
| ECE | 0.0211 |
Per-class F1:
| Label | Precision | Recall | F1 |
|---|---|---|---|
| Contradiction | 0.956 | 0.961 | 0.959 |
| Entailment | 0.962 | 0.955 | 0.959 |
| Neutral | 0.997 | 0.999 | 0.998 |
Limitations
- Domain specificity: Fine-tuned on CRS policy text with a specific hypothesis format (
{stakeholder}: {effect}). Performance on general NLI benchmarks (SNLI, MNLI, ANLI) is lower than the zero-shot base model — this is an expected cost of domain adaptation. - Hypothesis format: The model expects the stakeholder-prefixed hypothesis format. Inputs that deviate significantly from this format may produce unreliable probabilities.
- English only.
Citation
@misc{ashkinaze2026policyfx,
author = {Ashkinaze, Joshua},
title = {policyfx-nli: ModernBERT-NLI fine-tuned on Congressional Research Service policy reports},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/JoshuaAshkinaze/policyfx-nli}
}
- Downloads last month
- 15
Model tree for JoshuaAshkinaze/policyfx-nli
Base model
answerdotai/ModernBERT-base