MEDICAL CLASSIFICATION THROUGH MACHINE LEARNING, DOMAIN-SPECIFIC LANGUAGE MODELS AND GENERATIVE AI TECHNIQUES

dc.contributor.advisorHougen, Dean F
dc.contributor.authorGheibi, Reza
dc.contributor.committeeMemberBarnes, Ronald
dc.contributor.committeeMemberMudduluru, Sanjana
dc.contributor.committeeMemberPanjei, Egawati
dc.date.accessioned2026-07-06T19:10:32Z
dc.date.embargoExpiration
dc.date.issued2026
dc.date.proquestAvailable01/01/2026
dc.date.updated2026-07-06T19:10:32Z
dc.description.abstractAdvancing reliable machine learning and natural language processing methods for classification under data constrained conditions remains a central challenge in computer science. This dissertation addresses core problems related to data scarcity, class imbalance, heterogeneous tabular features, and the integration of structured and unstructured representations. To investigate these challenges, cervical cancer classification is used as an application domain where limited, noisy, and imbalanced datasets make model development particularly difficult. The work first evaluates traditional machine learning models including logistic regression, decision trees, random forests, gradient boosting, and multilayer perceptrons under varying data quality conditions. Multiple data augmentation strategies are examined, such as SMOTE, generative adversarial networks, diffusion based models, large language model, domain specific small language models, and synthetic tabular data generation, demonstrating that appropriate augmentation can substantially improve robustness and predictive performance. The second component introduces a framework that transforms structured medical records into natural language representations, allowing the use of modern language models for classification. Domain-specific and general purpose transformers (e.g., PubMedBERT, BioBERT, GPT-2, LLaMA, Mistral) are fine-tuned and evaluated in supervised, zero-shot, and few-shot settings. The results show that text-based modeling can match or exceed traditional ML performance, particularly in low-resource scenarios, while capturing richer contextual information. Finally, the dissertation extends this language-based methodology to a large-scale clinical taxonomy task. Through unsupervised analysis and systematic refinement, ambiguous specialty labels are consolidated into coherent Super-Specialties, and a fine-tuned PubMedBERT classifier achieves strong multi-label performance. This demonstrates the scalability of the proposed methods for organizing complex clinical data. In general, this dissertation demonstrates that combining structured data modeling, generative data augmentation, and language-based representations provides a flexible and effective framework for cervical cancer classification under data-constrained conditions. The findings highlight the potential of these approaches to support more accessible, scalable, and robust clinical decision-support systems, with broader implications for medical classification tasks beyond cervical cancer.
dc.identifier.urihttps://shareok.org//handle/11244/342760
dc.language.isoen
dc.publisherUniversity of Oklahoma – Graduate College
dc.subjectComputer science
dc.subjectInformation science
dc.subjectEngineering
dc.subjectComputer Science
dc.subjectMachine Learning
dc.subjectNatural Language Processing
dc.thesis.degreeD.Phil.
dc.titleMEDICAL CLASSIFICATION THROUGH MACHINE LEARNING, DOMAIN-SPECIFIC LANGUAGE MODELS AND GENERATIVE AI TECHNIQUES
ou.groupComputer Science: Engineering

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
Gheibi_oklahoma_2409A_10619.pdf
Size:
3.18 MB
Format:
Adobe Portable Document Format

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
2.01 KB
Format:
Item-specific license agreed upon to submission
Description: