Academic Direction

Research

Areas I am actively exploring and hope to contribute to as a graduate student and future researcher.

Research Statement

My research focuses on the reliability and robustness of AI systems in clinical decision-making. My M.Sc. thesis asks whether tool-using clinical LLMs correctly change their decisions when a small but clinically decisive variable crosses a meaningful decision boundary — studied through a counterfactual benchmark in clinical toxicology. More broadly, I work at the intersection of Clinical AI, reliable and robust language models, AI-enabled software systems, and the distributed and MLOps infrastructure that production AI depends on.

Current Research

M.Sc. Thesis · Artificial Intelligence
In Progress

Decision-Boundary Robustness in Tool-Using Clinical LLMs: A Counterfactual Benchmark in Clinical Toxicology

Islamic Azad University

My M.Sc. thesis investigates whether tool-using clinical LLMs correctly change their decisions when small but clinically decisive variables cross meaningful decision boundaries.

The work develops a structured counterfactual benchmark based on minimal case pairs in clinical toxicology. Each pair keeps the clinical context fixed while changing a single decision-relevant variable, allowing controlled evaluation of boundary recognition, unsafe use of invalid outputs, triage behavior, and correct next-step routing.

The study also evaluates runtime guards and structured decision reporting to better understand when LLM-based agents fail safely, fail silently, or correctly redirect the clinical workflow.

Research highlights

  • Counterfactual benchmark design
  • Tool-using clinical LLM agents
  • Decision-boundary robustness
  • Minimal-pair evaluation
  • Clinical toxicology
  • Structured, machine-readable evaluation
  • Runtime safety guards
  • Reproducible AI evaluation

Research Interests

Clinical AI & Decision Support

AI systems that support or evaluate clinical decision-making, with an emphasis on safety, validity, and clinically meaningful behavior.

Clinical Decision SupportClinical ToxicologySafetyClinical ValidityEvaluation

Reliable & Robust Large Language Models

Evaluation of LLM reliability, robustness, decision-boundary behavior, counterfactual sensitivity, and safe use of tool-augmented language models.

LLM EvaluationRobustnessDecision BoundariesCounterfactual SensitivityTool-Augmented LLMs

AI-Enabled Software Systems

Design and engineering of production software systems that integrate machine learning and generative AI into real-world workflows.

Generative AILLM IntegrationBackend SystemsReal-World Workflows

Distributed Systems, MLOps & Cloud Infrastructure

Reliable backend, distributed, and deployment infrastructure for AI applications, including reproducibility, asynchronous processing, containerization, and production ML systems.

Distributed SystemsMLOpsReproducibilityAsynchronous ProcessingContainerization

Background & Context

Industry Experience

My research is grounded in production work: I lead backend and DevOps development for a medical toxicology team, building AI-enabled software used by clinical toxicologists — which is where the thesis question comes from. Distributed-system failures and AI API orchestration in MedSpeech shape the systems side.

Academic Preparation

Undergraduate coursework in AI fundamentals (19.25/20), research methods (19/20), applied linear algebra (18.85/20), and software engineering, followed by graduate courses in Machine Learning and Deep Learning (17/20), provides the foundation for AI and systems research.

Current Trajectory

Pursuing an M.Sc. in Artificial Intelligence at Islamic Azad University with a thesis on decision-boundary robustness in tool-using clinical LLMs, while serving as Teaching Assistant for the Machine Learning course.

Long-Term Goal

Seeking a Ph.D. position in a group working on trustworthy AI agents, clinical AI evaluation, or reliable intelligent systems at scale.