Research
Areas I am actively exploring and hope to contribute to as a graduate student and future researcher.
Research Statement
My research focuses on the reliability and robustness of AI systems in clinical decision-making. My M.Sc. thesis asks whether tool-using clinical LLMs correctly change their decisions when a small but clinically decisive variable crosses a meaningful decision boundary — studied through a counterfactual benchmark in clinical toxicology. More broadly, I work at the intersection of Clinical AI, reliable and robust language models, AI-enabled software systems, and the distributed and MLOps infrastructure that production AI depends on.
Current Research
Decision-Boundary Robustness in Tool-Using Clinical LLMs: A Counterfactual Benchmark in Clinical Toxicology
Islamic Azad University
My M.Sc. thesis investigates whether tool-using clinical LLMs correctly change their decisions when small but clinically decisive variables cross meaningful decision boundaries.
The work develops a structured counterfactual benchmark based on minimal case pairs in clinical toxicology. Each pair keeps the clinical context fixed while changing a single decision-relevant variable, allowing controlled evaluation of boundary recognition, unsafe use of invalid outputs, triage behavior, and correct next-step routing.
The study also evaluates runtime guards and structured decision reporting to better understand when LLM-based agents fail safely, fail silently, or correctly redirect the clinical workflow.
Research highlights
- Counterfactual benchmark design
- Tool-using clinical LLM agents
- Decision-boundary robustness
- Minimal-pair evaluation
- Clinical toxicology
- Structured, machine-readable evaluation
- Runtime safety guards
- Reproducible AI evaluation
Research Interests
Clinical AI & Decision Support
AI systems that support or evaluate clinical decision-making, with an emphasis on safety, validity, and clinically meaningful behavior.
Reliable & Robust Large Language Models
Evaluation of LLM reliability, robustness, decision-boundary behavior, counterfactual sensitivity, and safe use of tool-augmented language models.
AI-Enabled Software Systems
Design and engineering of production software systems that integrate machine learning and generative AI into real-world workflows.
Distributed Systems, MLOps & Cloud Infrastructure
Reliable backend, distributed, and deployment infrastructure for AI applications, including reproducibility, asynchronous processing, containerization, and production ML systems.
Background & Context
Industry Experience
My research is grounded in production work: I lead backend and DevOps development for a medical toxicology team, building AI-enabled software used by clinical toxicologists — which is where the thesis question comes from. Distributed-system failures and AI API orchestration in MedSpeech shape the systems side.
Academic Preparation
Undergraduate coursework in AI fundamentals (19.25/20), research methods (19/20), applied linear algebra (18.85/20), and software engineering, followed by graduate courses in Machine Learning and Deep Learning (17/20), provides the foundation for AI and systems research.
Current Trajectory
Pursuing an M.Sc. in Artificial Intelligence at Islamic Azad University with a thesis on decision-boundary robustness in tool-using clinical LLMs, while serving as Teaching Assistant for the Machine Learning course.
Long-Term Goal
Seeking a Ph.D. position in a group working on trustworthy AI agents, clinical AI evaluation, or reliable intelligent systems at scale.