Instructions to use busum/biobert-lora-chameleon-radiology with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use busum/biobert-lora-chameleon-radiology with PEFT:
from peft import PeftModel from transformers import AutoModelForSequenceClassification base_model = AutoModelForSequenceClassification.from_pretrained("dmis-lab/biobert-v1.1") model = PeftModel.from_pretrained(base_model, "busum/biobert-lora-chameleon-radiology") - Transformers
How to use busum/biobert-lora-chameleon-radiology with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="busum/biobert-lora-chameleon-radiology")# Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("busum/biobert-lora-chameleon-radiology") model = AutoModel.from_pretrained("busum/biobert-lora-chameleon-radiology", device_map="auto") - Notebooks
- Google Colab
- Kaggle
BioBERT + LoRA: Chest CT Radiology Report Classifier (18 classes)
⚠️ Not for clinical use. This is an educational / portfolio project. It must not be used for diagnosis, triage or any decision about real patients.
A LoRA adapter for dmis-lab/biobert-v1.1 that classifies chest CT radiology reports into 18 classes (17 pathologies + Healthy). It was trained on synthetic reports from the Chameleon dataset; no real patient data was used.
Labels
Adenocarcinoma, Aspergilloma, Benign Nodules, Bronchiectasis, Bronchitis, Calcified Granulomas, Covid19, Emphysema, Healthy, Large Hodgkin Lymphoma, Metastatic Nodule, Non-Small Cell Lung Cancer, Pleural Effusion, Pneumonia (lobular, fungal pnemonia), Pulmonary Embolism (PE), Pulmonary Fibrosis, Sarcoidosis, Tuberculosis
(Label names are kept exactly as in the dataset, including the original typo "pnemonia".)
Preprocessing (important)
Reports were not fed to the model as-is. Before tokenization each report is reduced to
IMPRESSION: ... FINDINGS: ... (impression first), and the following are removed:
- clinical information, procedure, comparison and recommendation sections,
- "ATTENDING PHYSICIAN / RADIOLOGIST AGREEMENT" signature notes — their frequency differed strongly between classes (30%–66%), so they were a shortcut-learning risk,
- markdown
**markers.
This cut the median length from 514 to 350 BioBERT tokens; only 1% of reports exceed the 512-token limit, and because the impression comes first, truncation never removes it. Apply the same preprocessing at inference time, otherwise predictions silently degrade (training–serving skew).
Training
| Base model | dmis-lab/biobert-v1.1 @ 551ca18efd7f052c8dfa0b01c94c2a8e68bc5488 |
| Dataset | satvikt04/Chameleon-Radiology-Reports @ ebf03d97a699b6ad61fc20c9a93f08b54da13f06 |
| Split | 80 / 10 / 10 (8,000 / 1,000 / 1,000), stratified by label, seed 42 |
| LoRA | r=8, alpha=16, dropout=0.1, target modules query, value; classifier head fully trained |
| Trainable parameters | 308,754 of 108.6M (0.28%) |
| Optimisation | 3 epochs, lr 2e-4, batch size 16, fp16, max length 512 |
| Hardware | Google Colab, 1× T4 GPU, ~9.5 minutes |
| Libraries | transformers 5.18.0, peft 0.21.1, datasets 5.0.1 |
Validation:
| Epoch | Val loss | Accuracy | Macro-F1 |
|---|---|---|---|
| 1 | 0.229 | 0.983 | 0.967 |
| 2 | 0.030 | 0.996 | 0.992 |
| 3 | 0.023 | 0.996 | 0.992 |
Results (held-out test set, 1,000 reports)
| Accuracy | Macro-F1 |
|---|---|
| 0.995 | 0.991 |
Healthy: 500/500 correct. The few errors concentrate in clinically overlapping classes — lowest F1: Adenocarcinoma 0.947, Metastatic Nodule 0.951. Note that adenocarcinoma is itself a subtype of non-small cell lung cancer, so these labels are not mutually exclusive.
Limitations — read before trusting the numbers
- Synthetic, highly templated text. Train and test reports come from the same generator with near-identical phrasing, and the impression often names the pathology directly. The test score mostly measures how well the model learned this template, not real-world performance.
- Fails on out-of-distribution input. On three short sentences we wrote ourselves, the model was correct only once (Healthy, 0.975). "Moderate right pleural effusion…" was predicted as Bronchiectasis (0.23) and an explicit pulmonary embolism sentence as Covid19 (0.46). Low confidence is a useful warning signal — treat predictions below ~0.5 as unreliable.
- English chest CT reports only; single label per report; max 512 tokens.
- Not validated on any real clinical data.
How to use
import torch
from peft import PeftModel
from transformers import AutoConfig, AutoModelForSequenceClassification, AutoTokenizer
BASE = "dmis-lab/biobert-v1.1"
BASE_REVISION = "551ca18efd7f052c8dfa0b01c94c2a8e68bc5488"
ADAPTER = "busum/biobert-lora-chameleon-radiology"
config = AutoConfig.from_pretrained(ADAPTER) # contains the 18 label names
tokenizer = AutoTokenizer.from_pretrained(ADAPTER)
base = AutoModelForSequenceClassification.from_pretrained(BASE, revision=BASE_REVISION, config=config)
model = PeftModel.from_pretrained(base, ADAPTER).eval()
text = "IMPRESSION: No acute cardiopulmonary abnormality. FINDINGS: Lungs are clear." # apply the preprocessing above
enc = tokenizer(text, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
probs = model(**enc).logits.softmax(-1)[0]
top = probs.argmax().item()
print(config.id2label[top], round(probs[top].item(), 3))
The "classifier.weight / classifier.bias MISSING" message when loading the base model is expected: the adapter replaces the classifier with the trained one.
License and terms
- This adapter: CC BY-NC 4.0 (non-commercial).
- Training data: Chameleon dataset terms — research, educational and non-commercial use only; no redistribution of the data. The dataset itself is not included in this repository.
- Base model: BioBERT (dmis-lab), Apache 2.0.
Citation
Training data:
@misc{tripathi2025chameleon,
author = {Tripathi, Satvik and Enwerem, Don and Alkhulaifat, Dana and Sukumaran, Rithvik and Chambers, Charles and Levic, Darco and Daye, Dania and Cook, Tessa},
title = {Chameleon Dataset: A Large Multi-Pathology Synthetic Chest CT Radiology Reports Dataset},
year = {2025},
month = {July},
day = {10},
howpublished = {Available at SSRN},
url = {https://ssrn.com/abstract=5386019},
doi = {10.2139/ssrn.5386019}
}
Base model:
@article{lee2020biobert,
title = {BioBERT: a pre-trained biomedical language representation model for biomedical text mining},
author = {Lee, Jinhyuk and Yoon, Wonjin and Kim, Sungdong and Kim, Donghyeon and Kim, Sunkyu and So, Chan Ho and Kang, Jaewoo},
journal = {Bioinformatics},
volume = {36},
number = {4},
pages = {1234--1240},
year = {2020},
doi = {10.1093/bioinformatics/btz682}
}
- Downloads last month
- 10
Model tree for busum/biobert-lora-chameleon-radiology
Base model
dmis-lab/biobert-v1.1