BioBERT + LoRA: Chest CT Radiology Report Classifier (18 classes)

⚠️ Not for clinical use. This is an educational / portfolio project. It must not be used for diagnosis, triage or any decision about real patients.

A LoRA adapter for dmis-lab/biobert-v1.1 that classifies chest CT radiology reports into 18 classes (17 pathologies + Healthy). It was trained on synthetic reports from the Chameleon dataset; no real patient data was used.

Labels

Adenocarcinoma, Aspergilloma, Benign Nodules, Bronchiectasis, Bronchitis, Calcified Granulomas, Covid19, Emphysema, Healthy, Large Hodgkin Lymphoma, Metastatic Nodule, Non-Small Cell Lung Cancer, Pleural Effusion, Pneumonia (lobular, fungal pnemonia), Pulmonary Embolism (PE), Pulmonary Fibrosis, Sarcoidosis, Tuberculosis

(Label names are kept exactly as in the dataset, including the original typo "pnemonia".)

Preprocessing (important)

Reports were not fed to the model as-is. Before tokenization each report is reduced to IMPRESSION: ... FINDINGS: ... (impression first), and the following are removed:

  • clinical information, procedure, comparison and recommendation sections,
  • "ATTENDING PHYSICIAN / RADIOLOGIST AGREEMENT" signature notes — their frequency differed strongly between classes (30%–66%), so they were a shortcut-learning risk,
  • markdown ** markers.

This cut the median length from 514 to 350 BioBERT tokens; only 1% of reports exceed the 512-token limit, and because the impression comes first, truncation never removes it. Apply the same preprocessing at inference time, otherwise predictions silently degrade (training–serving skew).

Training

Base model dmis-lab/biobert-v1.1 @ 551ca18efd7f052c8dfa0b01c94c2a8e68bc5488
Dataset satvikt04/Chameleon-Radiology-Reports @ ebf03d97a699b6ad61fc20c9a93f08b54da13f06
Split 80 / 10 / 10 (8,000 / 1,000 / 1,000), stratified by label, seed 42
LoRA r=8, alpha=16, dropout=0.1, target modules query, value; classifier head fully trained
Trainable parameters 308,754 of 108.6M (0.28%)
Optimisation 3 epochs, lr 2e-4, batch size 16, fp16, max length 512
Hardware Google Colab, 1× T4 GPU, ~9.5 minutes
Libraries transformers 5.18.0, peft 0.21.1, datasets 5.0.1

Validation:

Epoch Val loss Accuracy Macro-F1
1 0.229 0.983 0.967
2 0.030 0.996 0.992
3 0.023 0.996 0.992

Results (held-out test set, 1,000 reports)

Accuracy Macro-F1
0.995 0.991

Healthy: 500/500 correct. The few errors concentrate in clinically overlapping classes — lowest F1: Adenocarcinoma 0.947, Metastatic Nodule 0.951. Note that adenocarcinoma is itself a subtype of non-small cell lung cancer, so these labels are not mutually exclusive.

Limitations — read before trusting the numbers

  • Synthetic, highly templated text. Train and test reports come from the same generator with near-identical phrasing, and the impression often names the pathology directly. The test score mostly measures how well the model learned this template, not real-world performance.
  • Fails on out-of-distribution input. On three short sentences we wrote ourselves, the model was correct only once (Healthy, 0.975). "Moderate right pleural effusion…" was predicted as Bronchiectasis (0.23) and an explicit pulmonary embolism sentence as Covid19 (0.46). Low confidence is a useful warning signal — treat predictions below ~0.5 as unreliable.
  • English chest CT reports only; single label per report; max 512 tokens.
  • Not validated on any real clinical data.

How to use

import torch
from peft import PeftModel
from transformers import AutoConfig, AutoModelForSequenceClassification, AutoTokenizer

BASE = "dmis-lab/biobert-v1.1"
BASE_REVISION = "551ca18efd7f052c8dfa0b01c94c2a8e68bc5488"
ADAPTER = "busum/biobert-lora-chameleon-radiology"

config = AutoConfig.from_pretrained(ADAPTER)  # contains the 18 label names
tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
base = AutoModelForSequenceClassification.from_pretrained(BASE, revision=BASE_REVISION, config=config)
model = PeftModel.from_pretrained(base, ADAPTER).eval()

text = "IMPRESSION: No acute cardiopulmonary abnormality. FINDINGS: Lungs are clear."  # apply the preprocessing above
enc = tokenizer(text, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    probs = model(**enc).logits.softmax(-1)[0]
top = probs.argmax().item()
print(config.id2label[top], round(probs[top].item(), 3))

The "classifier.weight / classifier.bias MISSING" message when loading the base model is expected: the adapter replaces the classifier with the trained one.

License and terms

  • This adapter: CC BY-NC 4.0 (non-commercial).
  • Training data: Chameleon dataset terms — research, educational and non-commercial use only; no redistribution of the data. The dataset itself is not included in this repository.
  • Base model: BioBERT (dmis-lab), Apache 2.0.

Citation

Training data:

@misc{tripathi2025chameleon,
  author = {Tripathi, Satvik and Enwerem, Don and Alkhulaifat, Dana and Sukumaran, Rithvik and Chambers, Charles and Levic, Darco and Daye, Dania and Cook, Tessa},
  title = {Chameleon Dataset: A Large Multi-Pathology Synthetic Chest CT Radiology Reports Dataset},
  year = {2025},
  month = {July},
  day = {10},
  howpublished = {Available at SSRN},
  url = {https://ssrn.com/abstract=5386019},
  doi = {10.2139/ssrn.5386019}
}

Base model:

@article{lee2020biobert,
  title = {BioBERT: a pre-trained biomedical language representation model for biomedical text mining},
  author = {Lee, Jinhyuk and Yoon, Wonjin and Kim, Sungdong and Kim, Donghyeon and Kim, Sunkyu and So, Chan Ho and Kang, Jaewoo},
  journal = {Bioinformatics},
  volume = {36},
  number = {4},
  pages = {1234--1240},
  year = {2020},
  doi = {10.1093/bioinformatics/btz682}
}
Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for busum/biobert-lora-chameleon-radiology

Adapter
(5)
this model

Dataset used to train busum/biobert-lora-chameleon-radiology