About
I’m a Ph.D. candidate in Biomedical Engineering at Johns Hopkins University, advised by Rama Chellappa. I combine biomedical expertise with research on language and vision-language models and multi-agent systems to advance AI capabilities and rigorously assess their reliability. Through hands-on system development and interdisciplinary collaboration, I study how inference-time computation, calibration, and agent interactions shape model performance, and how evaluation methods influence what we can conclude from it.
My collaborations with researchers at MIT and Microsoft span test-time scaling for clinical decision-making and hallucination-aware calibration for medical vision-language models. At Merck, I built a multi-agent system for scientific hypothesis generation and investigated how self-critique and evaluator choices affect assessments of its outputs. With ophthalmologists and entrepreneurs at the Wilmer Eye Institute, I worked on smartphone-based cataract screening evaluated prospectively in rural India. Across these projects, I translate domain-specific challenges into testable machine learning questions, carrying ideas through implementation, controlled experiments, and evaluation under practical constraints.
Experience
Graduate Research AssistantJohns Hopkins UniversityAug 2022 – Dec 2026
Towards Reliable Biomedical Decision-Making with Foundation Models.
- Hallucination-Aware Calibration: Developed vision-grounded calibration for medical VLMs, improving uncertainty estimates and error discrimination on open-ended questions, with ECE reductions up to 0.38 pp and AUROC gains up to 7.3 pp.COLM 2026
- Test-Time Scaling: Investigated inference-time sampling for clinical reasoning without fine-tuning, demonstrating AUC gains up to 30.4 pp across three diagnosis benchmarks and deriving analytical scaling laws.MIDL 2026
- Adaptive Inference: Developed adaptive token reduction and early exit for medical vision transformers, reducing FLOPs by 71.4% on average with only a 0.1 pp accuracy loss.MIDL 2026
- Adaptation under Distribution Shift: Developed cataract detection framework and led a comprehensive study of fine-tuning strategies under distribution shift and limited data; the resulting model achieved 91% accuracy in prospective deployment in rural India.NeurIPS 2024 AIM-FM Workshop
Research InternMerck & Co.Feb 2026 – Aug 2026
Multi-agent system for pharmaceutical toxicology hypothesis generation.
- Multi-Agent Hypothesis Generation: Built an end-to-end framework combining evidence retrieval, tool use, and iterative critique over proprietary multimodal data to generate mechanistic hypotheses for pharmaceutical toxicology.
- Agent Evaluation & LLM-as-a-Judge: Used controlled reruns and perturbation tests to evaluate self-critique and audit LLM judges, showing that initial critique drives most hypothesis changes and evaluation criteria shape measured diversity.NeurIPS 2026 AgenticLS Workshop
Research AssociateKAISTFeb 2019 – Apr 2022
Graph neural networks for Alzheimer's disease classification.
- GNN for Alzheimer's: bipartite subject–demographic graph with APPNP over resting-state fMRI and demographic data.ASMRM & ICMRI 2020Best Poster
- KAIST–KT Collaboration: managed a research partnership funded at $85K per year.
Education
2022 – 2026
Ph.D.
Biomedical Engineering
Johns Hopkins University
Advisor: Rama Chellappa
Thesis: Towards Reliable Biomedical Decision-Making with Foundation Models
2019 – 2021
M.S.
Bio and Brain Engineering
KAIST
Advisor: Yong Jeong
Thesis: Graph Neural Network for Predicting Alzheimer's Disease
2013 – 2018
B.S.
Bio and Brain Engineering
KAIST